Fourteen AI models tried to rebuild a photo in Blender
Three photos, fourteen models, up to 20 minutes and $4 per attempt. I rendered what each model saved and scored it. Astra still leads; the newest models are close behind, one of them for 36 cents an attempt.
PhotoAstra's scene · 67
PhotoAstra's scene · 70
PhotoAstra's scene · 62
Drag the divider. Left is the photo every model was given; right is what the selected model, GPT-6 Astra, built from it, rendered by the bench from the scene file it saved.
Score versus cost per attempt
Each dot is one model. Higher means a better score; farther left means a lower cost.
The job
For this experiment, an AI agent turns a photo into an editable Blender scene by writing and running code, rather than generating a mesh directly with a neural network. The agent follows an iterative process: it measures the picture, models the objects, sets the camera, lights the scene, renders, looks at its render next to the photo, and goes again. The result is a real Blender file with named, editable parts.
I am building a photo-to-Blender tool whose agent is powered by a language model. Until this test, I had used OpenAI's GPT-6 Astra, at about $4 a scene. That price is the reason for this test: if a cheaper model could do the job even nearly as well, the cost of a scene would fall several times over. So I built a bench where every model gets the same job under the same limits, and put fourteen on it.
The written brief asks for a faithful, editable reconstruction, forbids pasting the photo onto geometry, and tells the agent to save a usable scene early and then improve it, because the clock and the budget are enforced from outside by terminating it.
How I judged it
The tool normally scores the picture the agent hands in. For this benchmark, I instead open each delivered scene file and render it again at the reference photo’s aspect ratio, with a maximum dimension of 1,100 pixels and 96 samples. The scoring code then compares that render with the reference. A run that delivers no usable scene scores 0 and stays in the average.
The judge is deterministic image-and-mesh code, not an LLM: scorecard version 2, image metrics version 1.0-numpy, mesh audit version 1.0. Shape match uses estimated silhouette overlap (30%); detail uses edge F1, a measure of matching image edges (40%); colour uses multiscale RGB error (15%); and mesh health uses the weakest of five topology checks (15%). Those checks cover non-manifold and open edges, duplicate faces, inconsistent winding and degenerate triangles after position welding. Image comparisons use a maximum dimension of 512 pixels. The measurements are mapped through calibrated piecewise-linear scales to 0–100; these are comparison scores, not percentages of the scene recovered. All 36 delivered scenes have all four components. Download the run scores, costs and frozen board settings, and the same for the newest three.
Two more checks. A second render swings the camera 35° round the subject, so a flat stand-in behind the matched view would show. And every image file a scene uses is compared with the photo, since the brief forbids using it. No reference-file match was flagged in the delivered scenes, and I saw no obvious pasted-photo stand-in in the side views. Filename and digest checks can miss a renamed, re-encoded image; a single side view is also an inspection aid, not proof of complete 3D geometry.
The bench renders each scene itself. The picture a model hands in is never the one that gets scored.
The board
| # | Model | Score | Car | Desk | Shop | Cost / attempt | Attempts per $100 | Agent time | First render | Reasoning share |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 66 | 70 | 62 | 67 | $3.91 | 26 | 18 min | 4 min | 40% |
| 2 | GPT-6.1 SolOpenAI | 61 | 66 | 61 | 56 | $0.36 | 279 | 18 min | 3 min | – |
| 3 | Claude Opus 5.5Anthropic | 60 | 59 | 61 | 59 | $1.19 | 84 | 12 min | 3 min | – |
| 4 | Claude Sonnet 5.5Anthropic | 56 | 59 | 56 | 52 | $0.50 | 201 | 5 min | 4 min | – |
| 5 | GPT-5.6 SolOpenAI | 49 | 43 | 50 | 54 | $0.57 | 175 | 8 min | 2 min | 22% |
| 6 | Claude Fable 5.1Anthropic | 48 | 53 | 53 | 39 | $3.92 | 26 | 16 min | 5 min | 37% |
| 7 | ParetoUnbiased | 47 | 48 | 41 | 51 | $0.55 | 183 | 20 min | 4 min | 0% |
| 8 | Claude Opus 5Anthropic | 45 | 34 | 45 | 56 | $3.92 | 26 | 19 min | 5 min | 26% |
| 9 | GPT-5.6 TerraOpenAI | 44 | 39 | 39 | 55 | $0.33 | 300 | 8 min | 3 min | 13% |
| 10 | Kimi K3MoonshotAI | 41 | 46 | 33 | 44 | $1.55 | 65 | 17 min | 3 min | 12% |
| 11 | GLM 5.3 FlashZ.ai | 36 | 43 | 25 | 40 | $0.089 | 1,119 | 16 min | 3 min | 21% |
| 12 | Gemini 3.8 FlashGoogle | 34 | 40 | 28 | 34 | $2.10 | 48 | 16 min | 5 min | 24% |
| 13 | DeepSeek V4.1 FlashDeepSeek | 0 | – | – | – | $0.14 | 720 | 20 min | never | 83% |
| 14 | Qwen3.8 Max (0902)Qwen | 0 | – | – | – | $0.40 | 250 | 20 min | 12 min | 68% |
What they built
Numbers first, then the pictures. On each row the photo is outlined in green; each model's judged render follows with its score. Click any picture to see it large.
The car













Astra's car is boxier than the photo but unmistakably the same car, and it holds up from the side. Sol's is a toy from a different set; Fable's has melted at the front; Terra's is a flat slab with wheels; Gemini's is a loaf. Of the newest three, Sol 6.1's car is the closest to the photo, and Sonnet 5.5's is a chunkier toy on oversized wheels. Not shown: DeepSeek V4.1 Flash, Qwen3.8 Max (0902), which delivered no scene.
The desk













The simplest scene, and the one where Astra's lead is smallest: a mug, two books and a plant on a table. Astra and Fable are the two that get the mug's shape and the books' proportions. GLM's mug is a bucket and its plant a ring of leaves floating over the pot. The newest three all get the mug, the books and the plant; Opus 5.5 frames them smaller than the photo does. Not shown: DeepSeek V4.1 Flash, Qwen3.8 Max (0902), which delivered no scene.
The shopfront













The most complex scene. Astra's is the most complete, with the photo's stone, shutters and proportions all in place. The newest three all build a convincing corner building; Sonnet 5.5's is the flattest of them. Sol, Opus, Pareto and Terra built a corner building too, in flatter colours; Fable and Kimi flattened it into one wall; GLM's is a white box with a window. Not shown: DeepSeek V4.1 Flash, Qwen3.8 Max (0902), which delivered no scene.
Four things the board says
1. Astra works until the money runs out
Astra averaged 66 and was first on every photo. It was also stopped by the cost cap on every scene: it had a first test render 4 minutes in, then kept refining until, at about $3.90 and 18 minutes, the gateway refused the next request. It never ran out of time. The cost cap ended these three runs before the time limit. A larger budget would allow further attempts at refinement; whether those attempts improve the result needs another experiment.
2. The cheap OpenAI models stop early
GPT-5.6 Sol scored 49 for $0.57 a scene and GPT-5.6 Terra 44 for $0.33. Neither reached a time, cost or request cap. They declared the job finished after about 8 minutes, having spent $0.57 and $0.33 of a $4 budget on average. Their scenes are simpler than Astra's, not broken, and they had most of their budget and time left. That makes them a useful next experiment: a brief that asks for additional refinement within the remaining limits. If Sol moves from 49 towards 60 at a dollar a scene, the economics of the product change.
GPT-6.1 Sol, the newest of them, does not stop early. It worked for about 18 minutes a scene and scored 61 for 36 cents an attempt.
Sol and Terra stopped with substantial time and budget remaining. I have not yet tested whether more refinement would improve their scores.
3. Claude, and the cache
Claude Fable 5.1 scored 48 and Claude Opus 5 45, both at Astra's price. Fable's desk is the closest to Astra's on the board; its car melted and its shopfront is a pale single wall. Opus's shopfront (56) is its best scene and its car (34) its worst. The newest Claude models do much better: Claude Opus 5.5 scored 60 at $1.19 an attempt and Claude Sonnet 5.5 56 at 50 cents, the quickest on the board at about five minutes a scene.
Their first attempt did not count, and it is worth telling. In the gateway configuration used for the first attempts, Anthropic requests lacked the cache-control markers needed to enable caching. So on the first attempt 0% of their context was served from cache, against 96% for the OpenAI models. Each request incurred the full input charge for the conversation again. The archived Claude runs that reached the budget cap did so after eight to eighteen requests; Astra averaged about thirty-six. The gateway now adds the marker for Anthropic models, both were run again at 96–97% cached, and those are the numbers on the board. The archived uncached attempts logged about $16.39. They are excluded from the ranking but included in the experiment’s spending summary below.
4. The slow thinkers
DeepSeek V4.1 Flash and Qwen3.8 Max delivered nothing in 20 minutes, on any photo. Neither is incapable of the work. Both spend most of their output on private reasoning (83% and 68% of what they wrote), measuring the photo and fitting the camera at length, and the clock ran out before a scene was saved. DeepSeek made about as many requests as Astra in its 20 minutes; it just never rendered.
So I gave them more time, on separate boards so the results never mix with the 20-minute ones. With 60 minutes, DeepSeek built a lumpy car (44) and a one-wall shopfront (42) for about 60 cents each, and still failed the desk. With 120 minutes and the full $4, Qwen built a decent desk (51) and a car the score puts at 55 though it is a flattened body with a lump for a roof, and lost the shopfront when it called a web-search tool it had never been given. DeepSeek, given 120 minutes, delivered nothing at all, alongside a large slowdown in provider response speed, with individual responses taking 26 and 30 minutes. That makes it hard to separate model behaviour from serving conditions.




Given three and six times as long
What the two slow models built when the clock allowed it. Qwen used its whole $4 on each scene it finished, about Astra’s price for lower scores and roughly 97–100 minutes of agent time, versus Astra’s average of 18 minutes. The 120-minute allowance was six times the main board’s limit. DeepSeek's 60-minute scenes cost about 60 cents.
Neither is a candidate at the product's 20-minute limit. A brief that made them act before they had finished thinking might change that; the longer runs show that both models can deliver some of these scenes, but do not establish why the shorter attempts failed.
Also on the board
- Pareto (47, $0.55 estimated). This model ran anonymously as Union Alpha. It delivered every scene and sits with the older models in the 40s. OpenRouter charged $0 during the preview; repricing the logged tokens at Pareto’s revealed rates gives $0.55 per scene.
- Kimi K3 (41, $1.55). Busy and quick to a first render, and hit the request cap on two scenes. One run got an empty first answer from the model and ended after twelve seconds; the bench now restarts a run that ends within a minute with nothing saved, for every model alike, and that scene was run again.
- GLM 5.3 Flash (36, $0.089). Nine cents a scene, and it delivered all three. Crude, but its low cost makes it worth another look. Scores are a calibrated comparison scale, so dividing them by dollars is not a reliable measure of useful output.
- Gemini 3.8 Flash (34, $2.10). The Google endpoint used in this test rejected requests that ended on the model’s own message, which the agent tool sends whenever a model says what it will do next without calling a tool. Two attempts died on that before the gateway learned to add a closing turn for Google models. Its prompt was also only 47% cached, which contributed to its $2.10 average cost per attempt.
What I would not conclude
The newest three, GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5, ran later than the rest, on the current version of the brief and in their makers’ own agent CLIs. To see how much that changes, I ran GPT-6 Astra the same way: it scored 63 rather than 66. Read the gap between those three and the older models with that in mind.
There is one run per model per image. Small differences should be treated cautiously, and even Astra’s larger lead needs repeated runs to establish its stability. The brief was written and refined while Astra was the agent: every model got the same text, but it was shaped around one model’s habits. The OpenAI agent CLI also uses generic fallback settings for other models. Two providers needed gateway fixes before their three models could compete reliably. These are results for this model-plus-agent setup, not a model-neutral ranking or a proven lower bound on any model’s ability. Models running at the same time shared a GPU, so elapsed time also reflects contention.
The image components compare one rendered view with the reference; the mesh component checks a limited set of topology defects. Neither establishes the shape of hidden surfaces, physical dimensions, semantic part quality or ease of editing. Estimated silhouettes can be unreliable in full-frame photos, open surfaces may be intentional, and the mesh audit does not test self-intersections.
Cost, and what is next
The 42 accepted runs across the three boards logged $63.07: $50.79 for the main board, $1.68 for the 60-minute board and $10.60 for the 120-minute board. Archived attempts add about $17.64, including $16.39 for uncached Claude attempts, giving about $80.71 in recorded API spending. Some archived cost snapshots contain unsettled requests, so that is not a reconciled invoice total. Pareto’s preview was free; its estimated $1.64 replacement cost is excluded from these charged totals. Accepted runs alone cost about sixteen Astra attempts; including archived spending, about twenty-one. Local rendering and hardware costs are not included. The newest three ran on subscriptions; at list prices their nine attempts come to about $6.13, and the check run of Astra to about $5.08.
Next: all fourteen models on the current brief, three runs per image for the top group so that spread is visible, and a larger image set. Longer term, a model-neutral agent loop would remove the biggest asterisk on this board.
Method
Same job
The same three images, the same Blender installation, and a sandbox with no network or credentials available to the agent. The first eleven models shared one brief and the host gateway made their API calls; the newest three ran later on the current brief, in their makers’ own CLIs.
Same limits
20 minutes, $4 and 60 requests per scene. A model stopped by a cap is judged on what it had saved.
Judged on the benchmark render
Each scene file is rendered again by the server, at the reference photo's aspect ratio and 96 samples. That render is scored, never the model's own picture.
Seen from the side
A second render swings the camera 35° round, so a flat stand-in behind the matched view would show.
No pasted photo
Every image file a scene uses is compared with the reference by digest and by name.
Failures count
A run that delivers no scene scores 0 and stays in the average. A run that ends within a minute with nothing saved is started once more.
Real costs
Run costs come from the amounts reported through OpenRouter, not a separate invoice reconciliation. Pareto is one exception: its anonymous preview was free, and its displayed cost is reconstructed from logged token usage at the revealed rates. The newest three are the other: their costs are API-equivalent estimates from their CLIs’ token counts.
Equal plumbing
Where a provider needs something said before it behaves as others do unasked, the gateway says it: Anthropic's prompts are marked cacheable; Google gets a closing user turn. Each model's cache rate is recorded.
Frozen conditions
Caps, images, brief digest and scoring weights are stored with each board. The main board’s judging renderer was corrected to version 2 and its delivered scenes were rejudged together. The results here all use that renderer.
