Job Description
Anyone AI Labs is seeking a Research Scientist to own evaluation design for frontier models. You’ll design benchmarks across reasoning, coding, agents, tool-use and multi‑modal tasks, grounded in expert-verified truth and validated against multiple models.
You’ll QC results to survive buyer-side review and push for durable, informative evals as models improve. You’ll lead a team to recruit experts across coding and STEM, translate lab goals into robust evaluation pipelines, and publish public
#J-18808-LjbffrReady to Apply?
Take the next step in your AI career. Submit your application to Anyone AI today.
Submit Application