Coding agents can now turn a single photo into an editable 3D scene by writing Blender code, yet they cannot reliably judge whether each revision improves the result. The LEGO-Anything project from the University of Maryland and AWS pairs this Image-to-Code loop with LEGO-Bench, a test set of 208 images from 104 indoor and outdoor scenes. Its best tested model, GPT-6 Astra, reached 53.4 percent accuracy indoors and 39.6 percent outdoors. For business use, the format matters: an executable program can be inspected, edited and queried.

Coding agents turn photos into Blender scenes but fail at self-checks

How LEGO-Bench tests photo-to-code reconstruction

LEGO-Bench uses 443 registered assets and renders inputs from professionally built simulator scenes, which look natural while keeping exact geometry, depth and object assignments hidden as an answer key. Complexity can be increased without changing lighting or camera settings, avoiding the limits of real photos without ground truth and of synthetic scenes that look unrealistic. Scoring runs on three axes: validity checks whether a usable artifact was delivered, reconstruction measures visible geometry accuracy, and appearance compares a re-render pixel by pixel with the reference. All six tested GPT configurations delivered a working scene almost every time, while weaker setups scored around 15 percent.

The agent workflow is iterative rather than single-pass: it writes code, runs it, looks at the rendered result and revises toward the original image. Because the output explicitly captures objects, geometry, layout and camera position, the scene behaves like ordinary software that can be run, checked and modified. The researchers found that larger reasoning budgets helped GPT-6 variants substantially, with Astra rising from 32.3 to 61.8 percent on an office subset. Outdoor scenes proved harder than interiors, and accuracy fell as scene complexity grew. The design therefore trades one-shot speed for repeated correction cycles.

The background is a search for measurement that does not depend on an agent opinion. Analysis of work steps showed poor initial attempts, revisions that undid earlier progress, and unreliable self-assessment. When models chose which of two versions matched the original better, geometric judgments landed near or below chance level. The authors conclude refinement should rely on concrete measurements rather than agent judgment. That finding produced LEGO-Plugin, a no-training extension that anchors the starting scene in the reference image, replaces self-judgment with measurements and shields correct progress from regressive edits.

What photo-to-code means for product and operations teams

For companies working with catalogs, interiors, facilities or field sites, code-based scenes promise assets that designers and engineers can adjust directly instead of accepting a fixed mesh or image. Object detection, segmentation and depth estimation can be pulled directly from the program without extra training, which simplifies reuse across vision tasks. Object detection from reconstructions reached roughly half the performance of the specialized model DINO. Small teams gain a fast draft from one photo, while large organizations gain a structured artifact that fits review, versioning and downstream tooling.

The limitation is geometric fidelity and quality control. Segmentation and depth estimation lagged further behind specialized models such as SAM 3 and Depth Anything 3, and the authors describe current results as usable but unremarkable. LEGO-Plugin improved all six models, with weaker agents gaining up to 62.7 percent and the top model adding only about two percentage points, so gains depend on the starting point. This news alone does not mean single-photo reconstruction is ready for measurement, compliance or safety-critical geometry. Buyers should ask how validity, reconstruction and appearance are scored separately, how outdoor and complex scenes are handled, and what external checks prevent regressions.

A practical marker is whether measurement-guided refinement becomes standard in vendor demos and benchmarks, alongside scores split by indoor, outdoor and complexity levels. Related moves to watch include Unity plugins for Claude Code and Codex, direct reconstruction approaches such as Atlas from World Labs, and video-model routes such as GenCeption from Google DeepMind. If executable scenes close the gap to specialized vision models while keeping editability, photo-to-code will shift from rapid drafting to dependable operational data.