Camel turns a local model into a usability test—and ships 31 fixes
Camel’s second local-model benchmark improved stepwise task success from 81% to 92%, while exposing developer-experience failures headed for Camel 4.23.
Apache Camel is treating a small coding model less as a product feature than as an unusually fast usability tester. In the project’s second benchmark round, a 22 GB Qwen model worked through integration examples while a frontier model ran the harness and inspected failures; the resulting loop produced 33 issues, 31 of which Camel says are fixed and merged for version 4.23.
The benchmark moved closer to real work
The first round used isolated beginner examples. This round reorganized the test set into a ladder built around a fictional web shop, where orders, customers, a warehouse and a courier carry state from one exercise into the next. Camel also added a step-by-step mode: the harness starts an application with camel run --dev, asks the model to modify it incrementally through the Camel JBang MCP server, waits for reloads and then checks files and logs against the expected behavior, according to the engineering post.
That distinction matters because the published results improved most in the incremental test. The scored step-by-step set rose from 81% to 92% of steps passed, and fully successful runs increased from 29 of 50 to 41 of 50. One-shot attempts moved only from 66% to 70%; all 10 examples passed at least once across five attempts, but only four passed on every attempt. Camel also reports that tokens per one-shot attempt fell from about 3,100 to about 2,400.
The useful output is the fix list
The strongest result is not the model score. The benchmark exposed ordinary developer-experience defects: a bad edit during live reload could leave an application with no routes; adding another route file could trigger a duplicate route ID; newly created directories were not watched; and YAML load failures surfaced stack traces instead of the validator’s line-specific advice. The Camel team says those cases now preserve prior routes, improve file watching and route handling, and print more useful validation guidance.
The MCP surface also expanded. Camel says its server previously wrapped 27 of 53 runtime tools and omitted SQL, datasource, circuit-breaker, metrics, tracing and route-analysis operations; those gaps are now addressed for 4.23. File writes through the MCP server also wait for a reload and report whether it succeeded, failed or only refreshed properties, reducing the chance that an agent continues from a broken application state.
What integration teams should take from it
Camel’s own numbers are project-run measurements, not an independent comparison of coding models. The post is unusually candid about the remaining failures: valid applications still diverged from requested behavior, including merged routes, unexpected output shapes and PostgreSQL syntax sent to H2. In one experiment, giving the model SQL tools helped when the task explicitly asked for a database change, but did not make it proactively test an upsert before writing it.
The practical lesson is narrower and more useful than “AI writes integrations.” A repeatable agent harness can stress error messages, reload behavior and tool feedback at a pace that exposes friction human developers also meet. For Camel users, the concrete payoff is the 4.23 fix set; for platform teams building coding-agent workflows, the warning is that tool access helps with requested work but does not create verification discipline on its own.
sources
comments · 0