/Middle NLP Data Scientist

Middle NLP Data Scientist

PolandRemoteplvia direct
// Job Type
Full Time
// Salary
Not disclosed
// Posted
3 weeks ago
// Work Mode
remote

About the Role

Matching quality is at the core of the product. Behind it is an ML-powered ranking system that determines which jobs each candidate sees. We are looking for a Middle NLP Data Scientist to own and advance this ranking stack: fine-tuning transformer models, designing training data, and turning offline improvements into measurable production impact.

What You'll Do:

  • Design, fine-tune, and iterate on different NLP algorithms (bi-encoders, cross-encoders) for persona–job matching.
  • Work across the NLP model stack: bi-encoder embeddings, text-pair classification and cross-encoder reranking.
  • Build training data for ranking models: LLM-based labeling and improvement of human-annotated datasets.
  • Own offline and online evaluation: metric design (nDCG, recall@k), experimentation, and monitoring model performance in production.
  • Design feature pipelines from unstructured text (job descriptions, application forms, candidate profiles) and structured fields.
  • Build reusable, well-documented features suitable for both offline training and online inference.
  • Identify new signals and data sources that measurably improve matching quality.
  • Use LLMs for structured extraction from noisy text: prompt design for JSON output, schema validation, and automatic repair.
  • Audit existing data sources and build training-ready datasets for ML models.
  • Define data requirements with engineers: logging standards, formats, quality metrics, and monitoring signals.
  • Work closely with backend and ML engineers to bring models to production (batch and near–real-time), define inference APIs, and ensure observability.
  • Own the end-to-end lifecycle: from data and features, to models, to production impact.
  • Full understanding of the existing matching pipeline and its failure modes.
  • An upgrade to the ranking logic - strong improvement of current bi-encoders / cross-encoders.
  • A defined offline↔online evaluation loop and a clear baseline to measure every ranking change against.
  • Tangible, measured improvements in ranking quality (e.g. nDCG@10, recall@k) versus the current baseline.
  • 2+ years of experience in Data Science / ML with a focus on NLP.
  • Hands-on experience fine-tuning transformer models (BERT-family, sentence-transformers, cross-encoders / bi-encoders).
  • Solid understanding of text embeddings, semantic similarity, and reranking.
  • Experience building training datasets for NLP models (including synthetic / LLM-labeled data).
  • Strong Python skills and comfort with the PyTorch / HuggingFace ecosystem.
  • Experience taking ML models to production together with engineers.
  • Solid SQL knowledge and sound offline/online evaluation habits.
  • Experience with search, ranking, or recommender systems (retrieval, Learning-to-Rank, hybrid sparse+dense search, BM25 + embeddings).
  • Experience with vector databases / ANN indexes (Qdrant, FAISS, etc.) and search engines (Elasticsearch / OpenSearch).
  • Practical experience using LLMs in production (extraction, labelling, evaluation).
  • Familiarity with HR tech, ATS systems, or marketplace products.

What We Offer:

  • Market salary.
  • 20 workdays/year paid vacation.
  • Full-time remote work with flexible working hours.
  • Fast-paced, product-driven environment with real ownership and autonomy.
  • Opportunity to shape architecture and tech stack in a VC-backed startup from the early stages.
  • Close collaboration with a strong product team (PM, Design, Growth).
  • A culture of transparency, minimal bureaucracy, and quick decision-making.
  • Smart, ambitious teammates who value impact over process.

Dear Candidates, due to a high volume of applications, only selected candidates will be contacted for interviews. We appreciate your understanding. Thank you for considering a career with us.

  • Work type
  • Full-time
  • Location
  • Remote

Interested in this job?

Use our AI to tailor your resume for this Middle NLP Data Scientist position at hireforyou.pro.