Matching quality is at the core of the product. Behind it is an ML-powered ranking system that determines which jobs each candidate sees. We are looking for a Middle NLP Data Scientist to own and advance this ranking stack: fine-tuning transformer models, designing training data, and turning offline improvements into measurable production impact.
What You'll Do:
- Design, fine-tune, and iterate on different NLP algorithms (bi-encoders, cross-encoders) for persona–job matching.
- Work across the NLP model stack: bi-encoder embeddings, text-pair classification and cross-encoder reranking.
- Build training data for ranking models: LLM-based labeling and improvement of human-annotated datasets.
- Own offline and online evaluation: metric design (nDCG, recall@k), experimentation, and monitoring model performance in production.
- Design feature pipelines from unstructured text (job descriptions, application forms, candidate profiles) and structured fields.
- Build reusable, well-documented features suitable for both offline training and online inference.
- Identify new signals and data sources that measurably improve matching quality.
- Use LLMs for structured extraction from noisy text: prompt design for JSON output, schema validation, and automatic repair.
- Audit existing data sources and build training-ready datasets for ML models.
- Define data requirements with engineers: logging standards, formats, quality metrics, and monitoring signals.
- Work closely with backend and ML engineers to bring models to production (batch and near–real-time), define inference APIs, and ensure observability.
- Own the end-to-end lifecycle: from data and features, to models, to production impact.
- Full understanding of the existing matching pipeline and its failure modes.
- An upgrade to the ranking logic - strong improvement of current bi-encoders / cross-encoders.
- A defined offline↔online evaluation loop and a clear baseline to measure every ranking change against.
- Tangible, measured improvements in ranking quality (e.g. nDCG@10, recall@k) versus the current baseline.
- 2+ years of experience in Data Science / ML with a focus on NLP.
- Hands-on experience fine-tuning transformer models (BERT-family, sentence-transformers, cross-encoders / bi-encoders).
- Solid understanding of text embeddings, semantic similarity, and reranking.
- Experience building training datasets for NLP models (including synthetic / LLM-labeled data).
- Strong Python skills and comfort with the PyTorch / HuggingFace ecosystem.
- Experience taking ML models to production together with engineers.
- Solid SQL knowledge and sound offline/online evaluation habits.
- Experience with search, ranking, or recommender systems (retrieval, Learning-to-Rank, hybrid sparse+dense search, BM25 + embeddings).
- Experience with vector databases / ANN indexes (Qdrant, FAISS, etc.) and search engines (Elasticsearch / OpenSearch).
- Practical experience using LLMs in production (extraction, labelling, evaluation).
- Familiarity with HR tech, ATS systems, or marketplace products.
What We Offer:
- Market salary.
- 20 workdays/year paid vacation.
- Full-time remote work with flexible working hours.
- Fast-paced, product-driven environment with real ownership and autonomy.
- Opportunity to shape architecture and tech stack in a VC-backed startup from the early stages.
- Close collaboration with a strong product team (PM, Design, Growth).
- A culture of transparency, minimal bureaucracy, and quick decision-making.
- Smart, ambitious teammates who value impact over process.
Dear Candidates, due to a high volume of applications, only selected candidates will be contacted for interviews. We appreciate your understanding. Thank you for considering a career with us.
- Work type
- Full-time
- Location
- Remote