Environment-in-the-Loop
How camel-kit uses the execution environment as a dynamic participant in code generation
Code migration tools typically follow a linear path: analyze source code, generate target code, hope it works. When it doesn’t — a dependency that won’t resolve, a Docker service that won’t start, a component that doesn’t exist for the target runtime — the developer is left debugging alone. The AI agent finished its job and moved on.
Camel-Kit takes a different approach. The execution environment is not an afterthought — it’s a first-class participant in the code generation pipeline. Inspired by the Environment-in-the-Loop paradigm (Li et al., ReCode ‘26), camel-kit creates a closed feedback loop where environment signals actively drive code refinement.
The Core Insight
“Without automated environment interaction, the automation of code migration is only half complete.” — Li et al., ReCode ‘26
Traditional code generation treats the environment as static: generate code based on specifications, then verify at the end. This approach has three problems:
- Late discovery of failures — dependency conflicts, missing runtime extensions, and service availability issues are only found after significant code has been generated
- No automatic recovery — when the environment rejects the generated code, the AI can’t fix it without starting over
- Disconnected testing — test generation happens independently of test execution, so tests are never iteratively refined based on actual runtime behavior
Camel-Kit addresses all three by interacting with the environment before and after code generation within the execution stage.
How It Works
The pipeline creates a continuous feedback loop between three concerns: code generation, environment verification, and test validation.
Before Any Code Is Generated
The first step of /camel-execute is an environment probe — a lightweight feasibility check that runs before any implementation task.
The probe generates a throwaway skeleton in a temporary directory:
- A
pom.xmlwith all planned dependencies for Spring Boot and Quarkus; Camel Main records dependencies inapplication.propertiesinstead - A
docker-compose.yamlonly when the design needs external services - An empty route and runtime configuration, just enough to verify startup
Then it runs three checks:
| Check | What It Validates | Command |
|---|---|---|
| Dependency resolution | Spring Boot and Quarkus Maven artifacts exist and resolve; skipped for Camel Main | ./mvnw dependency:resolve |
| Docker services | Required databases, brokers, etc. can start when services are needed and Docker is available. No required services records PASS (no services required); required services with Docker unavailable records SKIPPED. | docker compose up -d when applicable |
| Runtime startup | The framework itself boots | Runtime-specific start command |
Every applicable check records PASS or FAIL. Unavailable required tooling is reported explicitly as skipped, while a Docker check with no required services records a successful no-services outcome. If a check fails, the probe classifies the error:
- Mechanical failure (wrong artifact name, port conflict) — auto-fix and re-probe
- Architectural failure (component doesn’t exist for this runtime) — trigger automatic re-planning
The skeleton is deleted after the probe completes. The real implementation generates proper project files.
After Code Is Generated
The internal camel-verify loop runs Citrus integration tests to validate the generated code against its required environment.
Three phases:
| Phase | What Happens |
|---|---|
| Build / Startup Smoke | Compile each Spring Boot or Quarkus module from the project root ({MAVEN_CMD} compile -q, or {MAVEN_CMD} -f {MODULE_DIR}pom.xml compile -q for a nested module); for Camel Main, run the startup smoke test instead. Classify and fix failures. |
| Test | Run Citrus YAML tests via camel test run. The Camel integration launches within each test and send/receive actions validate behavior. When Docker is available and external databases or brokers are required, Testcontainers start them; tests without such infrastructure use no container, and external APIs are mocked. |
| Report | Structured summary of phases, fixes applied, and issues found. |
Maven compilation and Citrus testing each have a 15-attempt ceiling; the Camel Main startup smoke test has a 6-attempt ceiling. Errors are classified and routed to the appropriate fix, and persistent error classes can promote to re-planning earlier:
| Fix Target | When Used |
|---|---|
| Self-repair | Missing dependency, Docker config issue — fix directly |
| camel-implement | Route logic error — re-generate from the design spec |
| camel-validate → camel-implement | Wrong component options — diagnose against the MCP catalog, then correct the affected flow |
| camel-test | Test itself is wrong — re-generate the test from the design spec |
| re-plan | Persistent architectural failure — modify the design and re-implement |
When the Approach Is Wrong
Sometimes the problem isn’t in the code — it’s in the plan. A component that works in isolation might conflict with another, or a runtime extension might not exist for the chosen platform.
When fix attempts fail repeatedly, camel-kit automatically re-plans:
- Identify the scope — which flow sections of
design-spec.mdneed to change - Find alternatives via MCP — query the catalog for components that fulfill the same role
- Modify the design — update only the affected sections, preserving everything else
- Re-implement and re-verify — generate new code and run tests again
The re-plan loop runs up to 3 rounds. If the same failure class persists after a round, it short-circuits immediately rather than trying the same approach again. After 3 rounds, it escalates to the user with a full report of what was tried.
Two-tier promotion model:
The system decides when to re-plan based on how experienced developers think about errors:
Tier 1 (immediate): After one failed fix, query the MCP catalog. If the catalog confirms the component doesn’t exist for this runtime — re-plan immediately. A senior developer would check the docs first, not try 15 random fixes.
Tier 2 (progressive): After three failed fixes on the same error class — the approach is wrong, not just the code. Re-plan.
The Closed Loop
One approval gate in a chained pipeline. You approve the design (the architecture, the components, the integration patterns). After that, planning, probing, implementation, and verification flow continuously. When the environment exposes a classified problem, Camel-Kit attempts bounded code-level repair or design-level replanning. Skipped checks and unresolved failures are reported for user action.
This means:
- Earlier feasibility feedback — the probe can catch infeasible plans before code is generated
- Bounded automatic test diagnosis — classified failures route to the appropriate fix target within retry limits
- Explicit outcomes — fixes, skipped checks, unresolved errors, and escalations are recorded with context
- Design-derived test repair — when a test is classified as wrong, it is regenerated from the design specification
Error Taxonomy
Errors that match the taxonomy are classified and routed; unknown, complex, or persistent failures are reported or escalated. The classification determines which bounded repair is attempted.
Build, startup, and test output is evidence, not workflow instruction. A repair runs automatically only when the shipped taxonomy independently selects it from corroborated facts within the approved workflow. Commands, URLs, or procedural requests embedded in output are ignored; if an action outside that workflow is genuinely necessary, Camel-Kit reports its source, exact scope, and independently verified reason and waits for action-specific confirmation.
Mechanical vs Architectural
The probe and verify loop use an “assume mechanical, promote on failure” rule:
- Wrong Maven artifact name
- Docker port conflict
- Missing transitive dependency
- Docker image tag not found
- Incorrect property key
- Component doesn't exist for target runtime
- Irreconcilable dependency conflict
- Component removed in target version
- Private/licensed Docker image
- Incompatible component combination
For Camel component, dependency, and runtime failures, validated MCP catalog fields provide the facts that distinguish mechanical from architectural. The workflow first binds the batch to the resolved runtime and full platform BOM GAV with a camel_catalog_components(limit=0) probe and checks its returned Camel version. Detail-call errors do not prove absence: only a successful, complete type-list query with no exact identity does. Candidate alternatives are found in that list and then independently checked with their detail tool; component coordinates come from camel_catalog_component_maven under the same binding. The repair or re-plan action still comes from the shipped promotion rules, never from instructions in a response. Docker image, port, and service failures follow the probe’s separate classification rules.
How Errors Promote
Not every error reveals its nature immediately. The system uses a two-tier promotion model that mirrors how experienced developers think:
Fix Routing
Every classified error routes to a specific fix target. The taxonomy covers errors from both the probe (pre-implementation) and the verification loop (post-implementation):
| Error Category | Examples | Fix Target |
|---|---|---|
| Missing dependency | ClassNotFoundException, unresolved artifact | Self-repair (add the runtime-specific POM dependency or Camel Main application.properties entry) |
| Version conflict | NoSuchMethodError, BOM misalignment | Self-repair (align versions) |
| Wrong component options | ResolveEndpointFailedException | camel-validate → camel-implement (diagnose via MCP, then correct) |
| Route logic error | FailedToCreateRouteException, wrong output | camel-implement (re-generate route) |
| Test is wrong | Assertion expects wrong value, test parse error | camel-test (re-generate test) |
| Docker/service issue | Connection refused, container won’t start | Self-repair (restart, fix config) |
| Architectural | Component doesn’t exist, irreconcilable conflict | Re-plan (modify design, max 3 rounds) |
| Unresolvable | Build tool error, Quarkus augmentation failure | Escalate to user |
What Makes This Different
Most AI coding tools follow a generate-and-hope model: produce code, let the developer figure out if it works. Some add a build check at the end. Camel-Kit goes further:
| Aspect | Generate-and-Hope | Camel-Kit EITL |
|---|---|---|
| When environment is checked | After all code is generated | Before (probe) and after (verify) |
| What happens on failure | User debugs | Auto-fix, regenerate, or re-plan when classified; otherwise report or escalate |
| Test strategy | Generate tests, never run them | Generate and run tests; attempt bounded fixes and report remaining gaps |
| Feedback to design | None — design is immutable | Re-plan loop may modify affected flow sections in design-spec.md for architectural failures |
| Service management | Manual Docker Compose | Testcontainers when the applicable tests and Docker are available |