Codesota · Tasks · Long-Horizon AutonomyHome/Tasks/Agents & Tool Use/Long-Horizon Autonomy

Long-Horizon Autonomy.

Completing tasks requiring many steps and delayed feedback.

0
Datasets
0
Results
Canonical metric
§ 02 · Canonical benchmark

The reference dataset.

Seeking canonical benchmark for this task.

Suggest one →
§ 03 · Top 10

Leading models.

Leading models across all datasets in this task.

No results yet. Be the first to contribute.

What were you looking for on Long-Horizon Autonomy?

Didn't find the model, metric, or dataset you needed? Tell us in one line. We read every message and reply within 48 hours.

§ 04 · All datasets

Tracked datasets.

0 datasets tracked for this task.

No datasets tracked yet.

§ 05 · Related tasks

Other tasks in Agents & Tool Use.

Tool CallingFunction CallingWeb AgentsDesktop AgentsBrowser AutomationComputer-Use AgentsResearch AgentsCustomer-Service Agents
Reply within 48 hours · No newsletter

Didn't find what you came for?

Still looking for something on Long-Horizon Autonomy? A missing model, a stale score, a benchmark we should cover — drop it here and we'll handle it.

Real humans read every message. We track what people are asking for and prioritize accordingly.