AI Software Engineer - Model Evaluation (f/m/d)
Benefits:
30 days of paid vacation
Access to a variety of fitness & wellness offerings via Wellhub
Mental health support through nilo.health
Substantially subsidized company pension plan
Subsidized Germany-wide transportation ticket
Education Requirements:
PhD in machine learning, NLP, statistics, or a related field (valued but not required)
Experience Requirements:
Experience with LLM evaluation, benchmark design, evaluation dataset curation, and experimental design
Familiarity with statistical methods for evaluation and experiment design
Track record of shipping impactful technical work (research or infrastructure)
Strong Python skills and comfort with ML tooling (PyTorch, evaluation frameworks, distributed systems)
Ability to reason about what an evaluation measures and whether it matters
Other Requirements:
Ownership mentality
Willingness to relocate to Heidelberg or travel regularly (potentially weekly)
Understanding of foundation model training (preferred)
German language proficiency (preferred)
Responsibilities:
Own benchmarks end-to-end: Select, implement, and maintain the evaluation suite used during pre-training
Build evaluation infrastructure: Develop and optimise the pipelines that run evaluations against training checkpoints
Design aggregation and reporting: Define how benchmark results translate into training decisions
Close capability gaps: Identify where models fall short and create benchmarks that measure progress
Own German evaluation: Ensure rigorous assessment of German language capabilities
Show more details