Codesota · RL EnvironmentsCode & software engineering← All environments
§ Ranked #01 by discriminative power

SWE-bench Verified.

An environment for Code & software engineering. Across 4 models with public results it spreads the best and worst 28%it still sorts frontier models, so training on it can still move yours.

Fallback tracked seed used by the CodeSOTA RL-environment views.

§ Public model scores

Who wins SWE-bench Verified.

Best public result per model entry, normalized 0..1. The spread between the top and bottom rows is what makes this environment worth — or not worth — a training run.

#Modelresolve rate
01Claude Sonnet72%
02OpenAI Codex69%
03Qwen Coder48%
04DeepSeek Coder44%
§ Nearby in the ranking
#EnvironmentSpreadDiscriminative
01SWE-bench VerifiedCode & software engineering28%0.28
02GAIAAgentic tool use26%0.26
03OSWorldComputer-use desktop and GUI25%0.25
§ Work with us

Need one that still separates models?

When the public environment for your capability saturates, you can’t tell your models apart and you can’t train past it. We build private, contamination-resistant, verifiable-reward environments and evals on a hold-out set — designed to discriminate where the public ones no longer do.