TolokaAI is building a dataset to evaluate AI coding agents by creating realistic developer tasks and an olympiad-style evaluation framework. You will design tasks, scripts, and criteria to assess agent performance within isolated development environments and real-world codebases.
Join a
remote-first, globally distributed team working on cutting-edge AI evaluation methods for top tech and research partners.
#J-18808-Ljbffr