micro1 is seeking skilled LLM Red-Teamers to join a high-impact project focused on the evaluation and improvement of frontier language models. You will play a critical role in shaping how next-generation AI systems learn and reason by providing high-quality, real-world input and rigorous testing.
Key responsibilities
- Develop complex, adversarial multi-turn conversations and task-based scenarios based on project specifications.
- Author precise evaluation rubrics to assess model responses against defined behavioral targets.
- Iteratively test conversations against frontier LLMs, escalating difficulty to reach quality thresholds.
- Deliver comprehensive task packages including transcripts, target behaviors, and supporting evidence.
- Validate LLM outputs by documenting model strengths and failure modes.
Requirements
- Exceptional written English skills with strong structural organization and clarity.
- Deep familiarity with large language models and the ability to identify common failure patterns.
- Proven critical thinking skills and meticulous attention to detail.
- Ability to work autonomously, interpreting complex specifications with minimal oversight.
- Background in writing-intensive or analysis-centric fields such as research, technical writing, or quality assurance.
What we offer
- Opportunity to contribute to the development of cutting-edge AI systems.
- Fully remote work environment allowing for independent contribution.
- Flexible, output-based task structure tailored to your workflow.