How we test models
Every night we run the same six tasks against each featured model, through the public AnyModl API, the way a coding agent calls it. Pass or fail is decided by code, not by another AI.
What a test looks like
Each request is OpenAI-compatible and shaped like an IDE agent's: streaming on, usage reporting on, tool
definitions sent with tool_choice: auto, and multi-turn tool loops where the model reads a file,
gets the result back, and decides what to do next. Calls go to https://api.anymodl.com/v1 with an
internal test key, so they hit the same routing and providers you do.
The six tasks (public set v1)
- Fix a bug: read a file, fix one function, write it back without touching the others.
- Edit a config: change one setting and keep every other line exactly as it was.
- Multi-step agent task: find an unknown file, read it, and create a new file from what it says (three or more tool calls).
- Long context: summarise about 30,000 tokens of a public-domain novel and find three codes hidden in it.
- Strict JSON: return an object with exactly the requested keys and types, no markdown.
- Plain explanation: explain a simple idea in two to four sentences.
Graders are deterministic: file contents are checked against patterns, JSON is parsed and type-checked, and answers are checked for required facts and length. A model that writes a file before reading it fails.
What the numbers mean
- Pass: share of the six tasks that passed tonight. Agent mode: share of the three tool-use tasks. “Tool calling tested” means every tool-use task passed and every tool call was valid.
- p90: 90th-percentile time per task, including every turn of a multi-step task.
- Cost per 1k passed tasks: what tonight's run cost at AnyModl customer prices, divided by tasks passed.
- Samples: how many tasks ran. Small samples move a lot from night to night, so we always show them.
What we don't claim
We test the request shape coding agents use, but we have not yet replayed requests captured from Cursor itself, so we say “tool calling tested”, not “works in Cursor”. A model that could not be reached tonight is left off the scorecard rather than shown as 0%. Results are a small nightly sample, not a benchmark leaderboard.
Dev1 AI's eval
The second table on the models page comes from Dev1 AI, our own assistant, which runs a separate nightly agent eval on its own work and routes requests on the results. Only aggregate numbers are published (pass rate, latency, cost, sample size), never prompts or task content.
Schedule: nightly at 04:15 UTC. Raw numbers: api.anymodl.com/v1/scores (JSON, no key needed).