Mindrift connects experienced software engineers with project-based opportunities focused on testing, evaluating, and improving frontier AI coding agents. As a Software Engineering Evaluation Specialist, you will design complex, realistic coding tasks that challenge AI systems to solve real-world development problems.
Key responsibilities
- Design realistic developer scenarios, such as bug fixes, broken ETL pipelines, or feature implementations.
- Build reproducible Docker environments with pinned dependencies to ensure consistent task execution.
- Write robust pytest suites that verify outcomes without leaking the solution to the AI agent.
- Create comprehensive instructions and reference solutions to ensure tasks are solvable and high-quality.
- Iterate on task designs based on expert feedback and eventually review tasks created by other authors.
Requirements
- 3+ years of professional software development experience in a backend stack like Python, Go, Node.js, Java, or Rust.
- Fluency in Python and pytest, including advanced features like fixtures, parametrization, and monkeypatching.
- Strong proficiency in Docker authoring and Linux/Bash for debugging containerized environments.
- Practical experience using AI coding agents on non-trivial development tasks.
- Professional proficiency in written English (B2+ level).
What we offer
- Flexible, project-based collaboration where you choose when and how to contribute.
- Opportunity to work on cutting-edge AI evaluation benchmarks.
- Competitive compensation based on task performance and expertise.
- A fully remote environment allowing you to work from anywhere.