Which AI designs the best parts?

nurb works with the AI subscription you already have. Every model gets the same real part-design jobs, and a machine grades the actual geometry against what was asked and against print physics.

// 01  the short answer

Start from what you subscribe to.

Have Claude?
run claude-fable-5
at medium effort
first-try prints18/18
time per part~3 min
$/part at API rates~$1.29
Have ChatGPT (Codex)?
run gpt-5.6-terra
at max effort
first-try prints30/30
time per part~5 min
$/part at API rates~$1.46
Have Grok?
run grok-4.6
at xhigh effort
first-try prints36/36
time per part~12 min
$/part at API rates~$0.27
// 02  the leaderboard

Every model, ranked.

Ranked by how often parts print right the first time. The six squares are the six jobs below, green to red; a dashed square is a job not yet run. Click a row for the per-attempt detail.

modelfirst-try printsjobstime$/part
1 claude-fable-5 · medium Claude 18/18 · 100% ~3 min ~$1.29

Eighteen attempts, eighteen parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. About three minutes and a dollar thirty a part, which makes this the quickest and cheapest clean sweep any Claude plan will give you: the same model at low effort also goes eighteen for eighteen but takes longer and costs more, and high effort wants two and a half times the money for the same result. Three attempts a job from one person, so a thinner sample than the Grok rows.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
2 claude-fable-5 · low Claude 18/18 · 100% ~4 min ~$1.78

Eighteen attempts, eighteen parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. Nothing is wrong with this row except the one above it: the same model at medium effort sweeps the same eighteen jobs faster and for fifty cents less a part, so there is no reason left to pick this one. Three attempts a job from one person, a thinner sample than the Grok rows.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
3 gpt-5.6-terra · max ChatGPT (Codex) 30/30 · 100% ~5 min ~$1.46

Thirty attempts, thirty parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. The first ChatGPT row to get everything right, at about five minutes and a dollar fifty a part. It reads the doctrine before it draws, and when a printability check complains it changes the shape rather than silencing the warning, which is the habit that matters most here. Two things to know before you run it. It guesses at command names often enough to break its own build about once every other part, though it always recovers within an edit or two. And it almost never chamfers, so parts come off it sharp, with no relief where they meet the bed: ask for the chamfers by name. Terra at xhigh gets twenty-nine of thirty a minute quicker for forty cents less, so what this row buys you is the last part.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
4 claude-fable-5 · high Claude 24/24 · 100% ~7 min ~$3.20

Twenty-four attempts, twenty-four parts worth printing, the largest clean sweep on the board. It checks its own work in every one of them: it cuts the part open, measures what it just built, and fixes what it finds before it stops. It is also the most expensive row here at over three dollars a part, and the same model at medium effort sweeps its own eighteen for a third of that, so pay this only for the extra checking.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
5 claude-opus-5 · high Claude 18/18 · 100% ~10 min ~$2.33

Eighteen attempts, eighteen parts worth printing, and honest about the unmeasured dimension every time. Around ten minutes a part, the slow end of the Claude rows, and the two design jobs are where that time goes. Opus at low effort misses one in twenty-four, runs in half the time and costs a third as much, so choose by whether you would rather wait once or re-run once. Opus at xhigh takes half again as long, costs a dollar more, and still misses one.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
6 grok-4.6 · xhigh Grok 36/36 · 100% ~12 min ~$0.27

Thirty-six attempts, thirty-six parts worth printing, pooled from two people's runs, and the only row on the board to sweep every job at this many attempts. It is also slow, about twelve minutes a part, where the same model at low effort takes two and a half and misses one in fifty-four. Pay the wait when the part matters; otherwise low is the Grok row.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
7 claude-fable-5 · max Claude 18/18 · 100% ~13 min ~$5.79

Eighteen attempts, eighteen parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. It checks its own work hard, better than twice the measuring, cutting open and re-rendering that the same model does at medium effort. That buys nothing here, because medium sweeps the same eighteen jobs in a quarter of the time for a quarter of the money. At about thirteen minutes and five dollars eighty a part, this is the most expensive row on the board, so pay it only for a part you cannot re-run. Three attempts a job from one person, so a thinner sample than the Grok rows.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
8 grok-4.6 · low Grok 53/54 · 98% ~2 min ~$0.071

Fifty-three of fifty-four right, pooled from two people's runs, at about two and a half minutes and seven cents a part. The curved pole rest and the D-shaft knob, the two jobs that catch most models, came out right on all nine attempts each. Its one miss was a wall clip you could not get a screwdriver into. Grok at xhigh is the only row that gets everything, but it takes five times as long for four times the money, so start here.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
96%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
9 grok-4.6 · medium Grok 47/48 · 98% ~7 min ~$0.18

Forty-seven of forty-eight right, pooled from three people's runs, and the one miss was a bit block that never got its top chamfer. Still the wrong Grok row to pick: low effort is on the same subscription and gets nearly the same share right in a third of the time for a third of the money, and xhigh gets everything. Skip it in both directions.

Follow the spec
100%
Survive the kernel
97%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
10 claude-opus-5 · xhigh Claude 29/30 · 97% ~15 min ~$3.20

Twenty-nine of thirty right, and honest about the unmeasured dimension every time. It checks its own work harder than any other row on the board: it runs the verification list on every single attempt and goes back to measure what it built in twenty-one. It asked for a picture of the part in twenty-nine of thirty and never got one, because the machine it ran on had no renderer installed, so every check it made was a measurement rather than a look. Its one miss belongs to the grader rather than the model, a knob whose six finger scoops came out the same diameter as the shaft hole, so the scorer drove the stem down a scoop; measured from the real bore the knob passes. At fifteen minutes and three dollars twenty a part it is the slowest Claude row here, and Opus at high effort sweeps its own eighteen in ten minutes for a dollar less.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
96%
Handle a missing measurement
100%
11 gpt-5.6-sol · max ChatGPT (Codex) 29/30 · 97% ~5 min ~$2.69

Twenty-nine of thirty right at about five minutes a part, and honest about the unmeasured dimension every time. It builds and rebuilds harder than any other Codex row, eight times a part at the middle, though it never renders what it made or runs the verification list. Its one miss is a wall clip whose screw hole it cut as a diamond so the hole would print without supports: the screw passes and the driver reaches, but the head lands on four corners instead of a full ring. Terra at max gets all thirty right in the same time for a dollar twenty less a part, so that is the row to run.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
92%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
12 gpt-5.6-terra · xhigh ChatGPT (Codex) 29/30 · 97% ~4 min ~$1.02

Twenty-nine of thirty right at about four minutes and a dollar a part, where the same model at high effort manages twenty-three. Effort is what terra was missing. It is the leanest worker here too, about twelve commands a part against twenty-six for the sol rows, and it was honest about the unmeasured dimension every time. Its one miss is a real one and a near one: a pole rest whose cradle it cut across only half the block, leaving the other half standing solid at exactly the pole height, so the pole could not come down into it. The clearance and the height were both right and the groove was one offset away. Terra at max gets that one too, for about a minute and forty-five cents more a part, and it is the row to pick now.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
90%
Make it fit
100%
Handle a missing measurement
100%
13 claude-opus-5 · low Claude 23/24 · 96% ~5 min ~$0.92

Twenty-three of twenty-four right at about five minutes and ninety cents a part, and honest about the unmeasured dimension every time. Its one miss was the easiest job on the board: a cable clip built to the stated size that stopped tracking once the size changed. Much the cheapest Opus row; high effort gets that last one right but takes twice as long for two and a half times the money.

Follow the spec
94%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
14 grok-4.6 · high Grok 20/21 · 95% ~11 min ~$0.24

Twenty of twenty-one right across all six jobs, and the one miss was a bit block missing its top chamfer. Still the wrong Grok row to pick: xhigh takes about the same time and gets everything right, and low effort runs five times faster for a third of the money.

Follow the spec
100%
Survive the kernel
92%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
15 claude-sonnet-5 · xhigh Claude 28/30 · 93% ~16 min ~$2.03

Twenty-eight of thirty right, and the slowest row on the board at about fifteen minutes a part. Five of the six jobs came out right on every attempt; the wall clip is the exception and it is where the time goes, averaging over half an hour with one attempt past forty minutes. Sonnet at high effort misses one more, runs four minutes quicker and costs fifty cents less a part, which is the better trade unless the part matters.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
88%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
16 gpt-5.6-sol · high ChatGPT (Codex) 28/30 · 93% ~4 min ~$1.76

Twenty-eight of thirty parts right at about four minutes each. Five of the six jobs came out right on every single attempt, the curved pole rest and the one-screw wall clip included, and both misses are the same mistake, a knob bored a shade too tight for the shaft to go in. The same model at medium effort gets the same twenty-eight right in less time for less money, and terra at max gets all thirty right for thirty cents less than this row.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
88%
Handle a missing measurement
100%
17 gpt-5.6-sol · medium ChatGPT (Codex) 28/30 · 93% ~3 min ~$1.63

Twenty-eight of thirty right at about three minutes a part, which is what the same model manages at high effort, sooner and for less. Four of the six jobs came out right every time, the curved pole rest among them. Its two misses were a wall clip that left the bundle nothing to sit against and a knob bored a shade too tight for the shaft. It was the ChatGPT row to pick until terra ran at max, which gets all thirty right for less than this one costs.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
93%
Design a curve
100%
Make it fit
94%
Handle a missing measurement
100%
18 claude-sonnet-5 · medium Claude 11/12 · 92% ~13 min ~$2.10

Eleven of twelve right, and the one miss is the easiest job here: a cable clip that stopped tracking its own dimensions once they changed. About ten minutes a part, and one wall clip attempt ran fifty minutes. Twelve attempts is half what the Sonnet rows around it carry, so read this as the thinnest Sonnet sample rather than the best one.

Follow the spec
74%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
19 claude-sonnet-5 · high Claude 27/30 · 90% ~12 min ~$1.58

Twenty-seven of thirty right, five of the six jobs perfect, and honest about the unmeasured dimension every time. All three misses are the same soft failure on the wall clip: it holds the bundle and takes its screw, it just spends more plastic than the job allowed. Around eleven minutes a part, and the wall clip is where that time goes.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
98%
Design a curve
100%
Make it fit
100%
Handle a missing measurement
100%
20 claude-opus-5 · max Claude 27/30 · 90% ~15 min ~$3.25

Twenty-seven of thirty right, and honest about the unmeasured dimension every time. Everything it made built, every one came out a single clean solid, and it never silenced a warning to get there. It also never saw any of them: the machine it ran on had no renderer, so it worked by measurement alone and closed every attempt with a table of what it had measured against what was asked. Its three misses hide in that table. Two leg cups took a 1.2mm chamfer on the corners, a size picked to clear a warning that a later change had already removed, and it cut into the wall the job wanted solid to the rim. One knob came out 26mm across the waist where the job asked for 28, because it measured its grip across a lobe instead of between them. All three would print, and all three are wrong somewhere the model was certain it had checked. Opus at xhigh gets one more right for the same time and the same money, and Opus at high sweeps its own eighteen in ten minutes for a dollar less, so there is no reason to pick this row.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
95%
Handle a missing measurement
91%
21 gpt-5.6-sol · xhigh ChatGPT (Codex) 27/30 · 90% ~4 min ~$2.11

Twenty-seven of thirty right at about four minutes a part, and the priciest Codex row on the board. Four of the six jobs came out right every time, the curved rest and the missing measurement among them. Its three misses look alike from the outside: the part passes every printability check and the interface is still wrong, two wall clips with no round bore for the screw and a knob whose socket opens on the bed instead of the top. Every other sol row gets more right for less, so there is no reason left to pick this one.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
84%
Design a curve
100%
Make it fit
94%
Handle a missing measurement
100%
22 gpt-5.6-luna · max ChatGPT (Codex) 32/36 · 89% ~5 min ~$0.13

Thirty-two of thirty-six right at thirteen cents a part, which is the most any ChatGPT row gets right for under a dollar. The wall clip, the cable clip, the curved rest and the missing measurement it got right on every one of six attempts. It slips on the bit block, where three of six wrote the pocket spacing down as a finished number instead of working it out, so the block comes out right at the size stated and loses its pockets once the bit grows, and once on a knob that read the shaft flat as twice its width and turned into a bore that spins. The same model at xhigh manages twenty-three of thirty for two cents less, so max is the luna row to run.

Follow the spec
100%
Survive the kernel
93%
Design from a problem
100%
Design a curve
100%
Make it fit
95%
Handle a missing measurement
100%
23 claude-opus-5 · medium Claude 16/18 · 89% ~7 min ~$1.66

Sixteen of eighteen right at about seven minutes a part, and honest about the unmeasured dimension every time. One knob came out too narrow to turn, barely half the grip width the job asked for, and one leg cup had walls that never reached the rim solid. Opus at low effort gets a better share right in less time for less money, and Opus at high effort gets everything, so this is the row to skip in both directions.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
100%
Design a curve
100%
Make it fit
92%
Handle a missing measurement
92%
24 gpt-5.6-sol · low ChatGPT (Codex) 15/18 · 83% ~2 min ~$1.62

Fifteen of eighteen right at about two minutes a part, which was the best Codex row until the same model ran at high effort. It got every stated dimension and the curved rest right every time. Its misses are about reaching the part rather than shaping it, two wall clips with no clear path in for the screw and its driver, and one knob bored too tight for the shaft. High effort costs about the same and misses less, so start there.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
73%
Design a curve
100%
Make it fit
90%
Handle a missing measurement
100%
25 claude-sonnet-5 · low Claude 10/12 · 83% ~6 min ~$0.84

Ten of twelve right, fast and cheap for a Claude plan, and it slips exactly where the jobs stop handing over dimensions: a wall clip with no way in for the screwdriver, and a rest the pole could not drop into. Fine for parts you spell out in full.

Follow the spec
100%
Survive the kernel
100%
Design from a problem
57%
Design a curve
83%
Make it fit
100%
Handle a missing measurement
100%
26 gpt-5.6-terra · medium ChatGPT (Codex) 37/48 · 77% ~3 min ~$0.84

Thirty-seven of forty-eight right across two people's runs, at about two and a half minutes a part. Every stated dimension and every missing measurement it handled right; the design jobs are where it thins out. Four of eight wall clips left no usable path for the screw and its driver, two pole rests came out too flat to cradle the pole, two knobs would not take the stem, and three bit blocks stopped tracking once the bit grew. The same model at xhigh gets twenty-nine of thirty for about twenty cents more, so spend the effort instead.

Follow the spec
100%
Survive the kernel
93%
Design from a problem
80%
Design a curve
89%
Make it fit
92%
Handle a missing measurement
100%
27 gpt-5.6-luna · xhigh ChatGPT (Codex) 23/30 · 77% ~5 min ~$0.12

Eleven cents a part, and twenty-three of thirty right where the same model at low effort managed five. Effort is what luna was missing. It still slips when a part has to keep working at other sizes: two bit blocks stopped building once the bit got bigger, and a wall clip put the screw through the only place the bundle had to sit. One knob came out with a round bore that spins on the shaft, and once it filed its guess at the unmeasured dimension without marking it as a guess.

Follow the spec
100%
Survive the kernel
95%
Design from a problem
96%
Design a curve
96%
Make it fit
94%
Handle a missing measurement
98%
28 gpt-5.6-terra · high ChatGPT (Codex) 23/30 · 77% ~3 min ~$0.79

Twenty-three of thirty right, and the wall clip is where it comes apart: four of five left no usable path for the screw and its driver. It also left one cable clip with no hole in its mounting tab, one knob the stem would not enter, and one bit block that stopped building once the bit grew. The curved rest and the missing measurement it got right every time. The same model at xhigh gets twenty-nine of thirty for about twenty cents more a part, which is the whole difference between these two rows.

Follow the spec
95%
Survive the kernel
97%
Design from a problem
69%
Design a curve
100%
Make it fit
94%
Handle a missing measurement
100%
29 gpt-5.6-luna · high ChatGPT (Codex) 22/30 · 73% ~4 min ~$0.10

Ten cents a part, and twenty-two of thirty right. The job it never got is the bit block: every one built at the size stated and then stopped tracking once the bit grew, losing its pockets or its top chamfer. Elsewhere it slipped once each, a wall clip with no way in for the screw, a rest the pole would not sit in at another size, and a knob with a round bore that spins on the shaft. The same model at xhigh is the same machine for two cents more.

Follow the spec
100%
Survive the kernel
84%
Design from a problem
92%
Design a curve
98%
Make it fit
94%
Handle a missing measurement
100%
30 gpt-5.6-luna · medium ChatGPT (Codex) 11/18 · 61% ~2 min ~$0.079

Eleven of eighteen right at about two minutes and a dime a part. All three knobs came out too tight for the stem, two wall clips stopped holding the bundle once its size changed, one bit block lost its chamfers when the bit grew, and once it wrote its guess at the unmeasured dimension down as though it had measured it. The same model at xhigh costs pennies more and misses less; run that.

Follow the spec
100%
Survive the kernel
96%
Design from a problem
93%
Design a curve
100%
Make it fit
71%
Handle a missing measurement
97%
31 gpt-5.6-terra · low ChatGPT (Codex) 10/18 · 56% ~2 min ~$0.74

Everything it made built, and about half were worth printing. The pattern is a part that works at the size you stated and nowhere else: all three of its wall clips stopped fitting when the cable bundle changed, and one pole rest came out flat where the job needed a curve. It never once went back to measure what it had made. The same model at xhigh is a different machine for about a quarter more a part.

Follow the spec
100%
Survive the kernel
78%
Design from a problem
68%
Design a curve
86%
Make it fit
90%
Handle a missing measurement
88%
32 gpt-5.6-luna · low ChatGPT (Codex) 5/18 · 28% ~3 min ~$0.077

Cheap, fast, and right five times out of eighteen. It wrote the pole's size straight into the file and still built a rest the pole would not drop into, at that size or any other. Elsewhere it left a 0.3mm wall no printer will lay down, and once wrote its guess at the unmeasured dimension down as though it had measured it. The same model at xhigh is a different machine; run that instead.

Follow the spec
91%
Survive the kernel
68%
Design from a problem
63%
Design a curve
63%
Make it fit
80%
Handle a missing measurement
84%
33 claude-haiku-4-5 · high Claude 2/12 · 17% ~8 min ~$0.46

The weakest row here, and the extra effort did not help. Two parts of twelve came out right. It also wrote its guess at the unmeasured dimension down as though it had measured it, which is the mistake nobody catches until the print is wrong six months later.

Follow the spec
100%
Survive the kernel
81%
Design from a problem
47%
Design a curve
47%
Make it fit
49%
Handle a missing measurement
50%
34 claude-haiku-4-5 · low Claude 4/33 · 12% ~7 min ~$0.46

Four parts of thirty-three came out right, the bottom of the board. The cable clip, the simplest job here, is the only job it got right more than once, and even the fully spelled-out bit block never once got its chamfers. All three design jobs failed on every attempt: poles that would not drop into the rest, knobs too tight for the stem, wall clips with no way in for the screw and its driver. Twice it wrote its guess at the unmeasured dimension down as though it had measured it. At around forty-six cents a part it costs six times what the best row on the board costs.

Follow the spec
85%
Survive the kernel
62%
Design from a problem
38%
Design a curve
47%
Make it fit
63%
Handle a missing measurement
80%
// 03  the tradeoff

Quality against speed.

Every dot is one model at one effort setting: how often its parts printed right, against how long it took per part. The best picks sit high and to the left, right most of the time without the wait.

0%25%50%75%100%36912151821minutes per part →printed right first tryclaude-fable-5MLgpt-5.6-terra+Hclaude-opus-5Hgrok-4.6X+LMXgpt-5.6-sol+XLHclaude-sonnet-5XHMMH+Xgpt-5.6-luna+MLLMXHHMLLclaude-haiku-4-5HL
ClaudeChatGPT (Codex)GrokL M H X +effort, low to max; hover a dot for its numbers
// 04  the jobs

Six jobs, graded on geometry.

Follow the specA cable clip with every dimension stated. Can it build exactly what you asked?
Survive the kernelA bit block dense with chamfers, right at the CAD kernel's limits. One wrong move and the part never builds.
Design from a problem“Hold this cable bundle on the wall with one screw.” No shape given: it has to design one that works and prints.
Design a curveA rest that must cradle a measured pole along a real arc. Flat answers touch at lines and lose; only curvature passes.
Make it fitA knob for a measured D-shaft. The grader drives the real stem: too tight jams, too loose rattles, a round bore spins.
Handle a missing measurementOne dimension nobody measured. Does it guess silently, or handle the unknown the honest way?
Add your model to this page. One line, your own subscription, a wizard for the rest. Every run pools with everyone else's. $curl -fsSL https://nurb.dev/bench.sh | sh Or paste that line to your AI and let it drive.

$/part is what the same tokens would cost at API list prices; on a subscription it comes out of your plan.

Early days: 912 graded parts across 6 jobs so far. Each bar averages every attempt on file, and the ticks are the attempts themselves.

Grading is a fixed rubric measured on the part's actual geometry, so the only randomness is the model's. Raw results, full transcripts, and the grading code are on GitHub.