Coding & agent model shortlist
Compare deployment candidates for coding and tool-using agents, with model-card sources, runtime questions, and an application-level evaluation plan.
- What is included
- Selected profiles whose publisher descriptions cover coding or agent workloads. Model sizes vary substantially; inclusion does not mean equivalent hardware needs or verified tool compatibility.
- How this list is ordered
- Alphabetical by model name. No benchmark score or performance order is assigned.
Publisher sources are linked for every model. No independent performance scores have been measured. On smaller screens, scroll the table horizontally.
| Model | Architecture | Context | Deployment note |
|---|---|---|---|
| MoE · multimodal | 1M | A large sparse checkpoint for testing multi-step tool workflows.Compare model | |
| MoE · vision | 256K | Reasoning and coding candidate; Small is a product name, not a memory guarantee.Compare model | |
| Hybrid MoE | 1M | An official quantized agent candidate; validate the NVFP4 deployment path.Compare model | |
| Hybrid MoE | 256K | Coding-focused; evaluate the actual tool parser and agent scaffold.Compare model | |
| Dense · vision | 256K native | A dense candidate for a different hardware budget from the large sparse entries.Compare model |
Rank completed tasks when measurements arrive
For a coding agent, generating tokens is only one part of the job. Record whether the change passes its tests, whether tool calls are valid, and how many attempts are required. Keep the repository snapshot, agent instructions, tool access, and task budget fixed so a future ranking measures comparable work.
The serving configuration includes the scaffold
Chat templates, tool-call parsing, reasoning settings, and maximum output length affect the behavior an agent sees. Validate one complete tool cycle before a longer evaluation. A successful plain-text chat request does not prove that a checkpoint works with a particular coding harness.
Separate quality from capacity
First reject configurations that cannot complete the target tasks reliably. Then compare end-to-end time and completed tasks within a fixed resource budget. No independent coding scores or deployment winners have been measured by BenchGrid yet; these entries are candidates for that experiment.
Before you choose hardware
A good benchmark shows its working. explains the next planning step. For the distinction between publisher specifications, calculations, and measurements, read the BenchGrid methodology.