micro1 is seeking a Member of Technical Staff to join our team and advance the evaluation and development of frontier coding agents. This role sits at the intersection of AI research, software engineering, and model evaluation, where you will design the benchmarks, methodologies, and data systems that shape how next-generation coding models are measured and improved.
Key responsibilities
- Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methodologies, and quality standards.
- Lead end-to-end research initiatives focused on measuring and improving coding model performance across diverse software engineering tasks.
- Develop high-quality datasets, golden examples, and evaluation protocols that enable reliable assessment of frontier coding systems.
- Analyze model behavior and failure modes, identifying systematic weaknesses and translating findings into actionable improvements.
- Build tooling and infrastructure that support large-scale experimentation, data generation, and evaluation pipelines.
Requirements
- Strong software engineering background with expertise in Python, C++, or comparable programming languages.
- 3+ years of experience in software engineering, machine learning, AI research, evaluation, or related technical disciplines.
- Experience designing, reviewing, or validating technical assessments, benchmarks, or evaluation methodologies.
- Familiarity with large language models, coding agents, reinforcement learning, or related AI systems.
- Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
What we offer
- Opportunity to work on frontier AI systems and shape the future of coding agents.
- Collaborative environment working closely with researchers, engineers, and applied AI teams.
- Support for professional growth in a fast-moving research environment.
- Comprehensive benefits package designed to support a high-performing, remote-first workforce.