Which AI designs the best parts?
nurb works with the AI subscription you already have. Every model gets the same real part-design jobs, and a machine grades the actual geometry against what was asked and against print physics.
Start from what you subscribe to.
Every model, ranked.
Ranked by how often parts print right the first time. The six squares are the six jobs below, green to red; a dashed square is a job not yet run. Click a row for the per-attempt detail.
1 claude-fable-5 · medium Claude 18/18 · 100% ~3 min ~$1.29
Eighteen attempts, eighteen parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. About three minutes and a dollar thirty a part, which makes this the quickest and cheapest clean sweep any Claude plan will give you: the same model at low effort also goes eighteen for eighteen but takes longer and costs more, and high effort wants two and a half times the money for the same result. Three attempts a job from one person, so a thinner sample than the Grok rows.
2 claude-fable-5 · low Claude 18/18 · 100% ~4 min ~$1.78
Eighteen attempts, eighteen parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. Nothing is wrong with this row except the one above it: the same model at medium effort sweeps the same eighteen jobs faster and for fifty cents less a part, so there is no reason left to pick this one. Three attempts a job from one person, a thinner sample than the Grok rows.
3 gpt-5.6-terra · max ChatGPT (Codex) 30/30 · 100% ~5 min ~$1.46
Thirty attempts, thirty parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. The first ChatGPT row to get everything right, at about five minutes and a dollar fifty a part. It reads the doctrine before it draws, and when a printability check complains it changes the shape rather than silencing the warning, which is the habit that matters most here. Two things to know before you run it. It guesses at command names often enough to break its own build about once every other part, though it always recovers within an edit or two. And it almost never chamfers, so parts come off it sharp, with no relief where they meet the bed: ask for the chamfers by name. Terra at xhigh gets twenty-nine of thirty a minute quicker for forty cents less, so what this row buys you is the last part.
4 claude-fable-5 · high Claude 24/24 · 100% ~7 min ~$3.20
Twenty-four attempts, twenty-four parts worth printing, the largest clean sweep on the board. It checks its own work in every one of them: it cuts the part open, measures what it just built, and fixes what it finds before it stops. It is also the most expensive row here at over three dollars a part, and the same model at medium effort sweeps its own eighteen for a third of that, so pay this only for the extra checking.
5 claude-opus-5 · high Claude 18/18 · 100% ~10 min ~$2.33
Eighteen attempts, eighteen parts worth printing, and honest about the unmeasured dimension every time. Around ten minutes a part, the slow end of the Claude rows, and the two design jobs are where that time goes. Opus at low effort misses one in twenty-four, runs in half the time and costs a third as much, so choose by whether you would rather wait once or re-run once. Opus at xhigh takes half again as long, costs a dollar more, and still misses one.
6 grok-4.6 · xhigh Grok 36/36 · 100% ~12 min ~$0.27
Thirty-six attempts, thirty-six parts worth printing, pooled from two people's runs, and the only row on the board to sweep every job at this many attempts. It is also slow, about twelve minutes a part, where the same model at low effort takes two and a half and misses one in fifty-four. Pay the wait when the part matters; otherwise low is the Grok row.
7 claude-fable-5 · max Claude 18/18 · 100% ~13 min ~$5.79
Eighteen attempts, eighteen parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. It checks its own work hard, better than twice the measuring, cutting open and re-rendering that the same model does at medium effort. That buys nothing here, because medium sweeps the same eighteen jobs in a quarter of the time for a quarter of the money. At about thirteen minutes and five dollars eighty a part, this is the most expensive row on the board, so pay it only for a part you cannot re-run. Three attempts a job from one person, so a thinner sample than the Grok rows.
8 grok-4.6 · low Grok 53/54 · 98% ~2 min ~$0.071
Fifty-three of fifty-four right, pooled from two people's runs, at about two and a half minutes and seven cents a part. The curved pole rest and the D-shaft knob, the two jobs that catch most models, came out right on all nine attempts each. Its one miss was a wall clip you could not get a screwdriver into. Grok at xhigh is the only row that gets everything, but it takes five times as long for four times the money, so start here.
9 grok-4.6 · medium Grok 47/48 · 98% ~7 min ~$0.18
Forty-seven of forty-eight right, pooled from three people's runs, and the one miss was a bit block that never got its top chamfer. Still the wrong Grok row to pick: low effort is on the same subscription and gets nearly the same share right in a third of the time for a third of the money, and xhigh gets everything. Skip it in both directions.
10 claude-opus-5 · xhigh Claude 29/30 · 97% ~15 min ~$3.20
Twenty-nine of thirty right, and honest about the unmeasured dimension every time. It checks its own work harder than any other row on the board: it runs the verification list on every single attempt and goes back to measure what it built in twenty-one. It asked for a picture of the part in twenty-nine of thirty and never got one, because the machine it ran on had no renderer installed, so every check it made was a measurement rather than a look. Its one miss belongs to the grader rather than the model, a knob whose six finger scoops came out the same diameter as the shaft hole, so the scorer drove the stem down a scoop; measured from the real bore the knob passes. At fifteen minutes and three dollars twenty a part it is the slowest Claude row here, and Opus at high effort sweeps its own eighteen in ten minutes for a dollar less.
11 gpt-5.6-sol · max ChatGPT (Codex) 29/30 · 97% ~5 min ~$2.69
Twenty-nine of thirty right at about five minutes a part, and honest about the unmeasured dimension every time. It builds and rebuilds harder than any other Codex row, eight times a part at the middle, though it never renders what it made or runs the verification list. Its one miss is a wall clip whose screw hole it cut as a diamond so the hole would print without supports: the screw passes and the driver reaches, but the head lands on four corners instead of a full ring. Terra at max gets all thirty right in the same time for a dollar twenty less a part, so that is the row to run.
12 gpt-5.6-terra · xhigh ChatGPT (Codex) 29/30 · 97% ~4 min ~$1.02
Twenty-nine of thirty right at about four minutes and a dollar a part, where the same model at high effort manages twenty-three. Effort is what terra was missing. It is the leanest worker here too, about twelve commands a part against twenty-six for the sol rows, and it was honest about the unmeasured dimension every time. Its one miss is a real one and a near one: a pole rest whose cradle it cut across only half the block, leaving the other half standing solid at exactly the pole height, so the pole could not come down into it. The clearance and the height were both right and the groove was one offset away. Terra at max gets that one too, for about a minute and forty-five cents more a part, and it is the row to pick now.
13 claude-opus-5 · low Claude 23/24 · 96% ~5 min ~$0.92
Twenty-three of twenty-four right at about five minutes and ninety cents a part, and honest about the unmeasured dimension every time. Its one miss was the easiest job on the board: a cable clip built to the stated size that stopped tracking once the size changed. Much the cheapest Opus row; high effort gets that last one right but takes twice as long for two and a half times the money.
14 grok-4.6 · high Grok 20/21 · 95% ~11 min ~$0.24
Twenty of twenty-one right across all six jobs, and the one miss was a bit block missing its top chamfer. Still the wrong Grok row to pick: xhigh takes about the same time and gets everything right, and low effort runs five times faster for a third of the money.
15 claude-sonnet-5 · xhigh Claude 28/30 · 93% ~16 min ~$2.03
Twenty-eight of thirty right, and the slowest row on the board at about fifteen minutes a part. Five of the six jobs came out right on every attempt; the wall clip is the exception and it is where the time goes, averaging over half an hour with one attempt past forty minutes. Sonnet at high effort misses one more, runs four minutes quicker and costs fifty cents less a part, which is the better trade unless the part matters.
16 gpt-5.6-sol · high ChatGPT (Codex) 28/30 · 93% ~4 min ~$1.76
Twenty-eight of thirty parts right at about four minutes each. Five of the six jobs came out right on every single attempt, the curved pole rest and the one-screw wall clip included, and both misses are the same mistake, a knob bored a shade too tight for the shaft to go in. The same model at medium effort gets the same twenty-eight right in less time for less money, and terra at max gets all thirty right for thirty cents less than this row.
17 gpt-5.6-sol · medium ChatGPT (Codex) 28/30 · 93% ~3 min ~$1.63
Twenty-eight of thirty right at about three minutes a part, which is what the same model manages at high effort, sooner and for less. Four of the six jobs came out right every time, the curved pole rest among them. Its two misses were a wall clip that left the bundle nothing to sit against and a knob bored a shade too tight for the shaft. It was the ChatGPT row to pick until terra ran at max, which gets all thirty right for less than this one costs.
18 claude-sonnet-5 · medium Claude 11/12 · 92% ~13 min ~$2.10
Eleven of twelve right, and the one miss is the easiest job here: a cable clip that stopped tracking its own dimensions once they changed. About ten minutes a part, and one wall clip attempt ran fifty minutes. Twelve attempts is half what the Sonnet rows around it carry, so read this as the thinnest Sonnet sample rather than the best one.
19 claude-sonnet-5 · high Claude 27/30 · 90% ~12 min ~$1.58
Twenty-seven of thirty right, five of the six jobs perfect, and honest about the unmeasured dimension every time. All three misses are the same soft failure on the wall clip: it holds the bundle and takes its screw, it just spends more plastic than the job allowed. Around eleven minutes a part, and the wall clip is where that time goes.
20 claude-opus-5 · max Claude 27/30 · 90% ~15 min ~$3.25
Twenty-seven of thirty right, and honest about the unmeasured dimension every time. Everything it made built, every one came out a single clean solid, and it never silenced a warning to get there. It also never saw any of them: the machine it ran on had no renderer, so it worked by measurement alone and closed every attempt with a table of what it had measured against what was asked. Its three misses hide in that table. Two leg cups took a 1.2mm chamfer on the corners, a size picked to clear a warning that a later change had already removed, and it cut into the wall the job wanted solid to the rim. One knob came out 26mm across the waist where the job asked for 28, because it measured its grip across a lobe instead of between them. All three would print, and all three are wrong somewhere the model was certain it had checked. Opus at xhigh gets one more right for the same time and the same money, and Opus at high sweeps its own eighteen in ten minutes for a dollar less, so there is no reason to pick this row.
21 gpt-5.6-sol · xhigh ChatGPT (Codex) 27/30 · 90% ~4 min ~$2.11
Twenty-seven of thirty right at about four minutes a part, and the priciest Codex row on the board. Four of the six jobs came out right every time, the curved rest and the missing measurement among them. Its three misses look alike from the outside: the part passes every printability check and the interface is still wrong, two wall clips with no round bore for the screw and a knob whose socket opens on the bed instead of the top. Every other sol row gets more right for less, so there is no reason left to pick this one.
22 gpt-5.6-luna · max ChatGPT (Codex) 32/36 · 89% ~5 min ~$0.13
Thirty-two of thirty-six right at thirteen cents a part, which is the most any ChatGPT row gets right for under a dollar. The wall clip, the cable clip, the curved rest and the missing measurement it got right on every one of six attempts. It slips on the bit block, where three of six wrote the pocket spacing down as a finished number instead of working it out, so the block comes out right at the size stated and loses its pockets once the bit grows, and once on a knob that read the shaft flat as twice its width and turned into a bore that spins. The same model at xhigh manages twenty-three of thirty for two cents less, so max is the luna row to run.
23 claude-opus-5 · medium Claude 16/18 · 89% ~7 min ~$1.66
Sixteen of eighteen right at about seven minutes a part, and honest about the unmeasured dimension every time. One knob came out too narrow to turn, barely half the grip width the job asked for, and one leg cup had walls that never reached the rim solid. Opus at low effort gets a better share right in less time for less money, and Opus at high effort gets everything, so this is the row to skip in both directions.
24 gpt-5.6-sol · low ChatGPT (Codex) 15/18 · 83% ~2 min ~$1.62
Fifteen of eighteen right at about two minutes a part, which was the best Codex row until the same model ran at high effort. It got every stated dimension and the curved rest right every time. Its misses are about reaching the part rather than shaping it, two wall clips with no clear path in for the screw and its driver, and one knob bored too tight for the shaft. High effort costs about the same and misses less, so start there.
25 claude-sonnet-5 · low Claude 10/12 · 83% ~6 min ~$0.84
Ten of twelve right, fast and cheap for a Claude plan, and it slips exactly where the jobs stop handing over dimensions: a wall clip with no way in for the screwdriver, and a rest the pole could not drop into. Fine for parts you spell out in full.
26 gpt-5.6-terra · medium ChatGPT (Codex) 37/48 · 77% ~3 min ~$0.84
Thirty-seven of forty-eight right across two people's runs, at about two and a half minutes a part. Every stated dimension and every missing measurement it handled right; the design jobs are where it thins out. Four of eight wall clips left no usable path for the screw and its driver, two pole rests came out too flat to cradle the pole, two knobs would not take the stem, and three bit blocks stopped tracking once the bit grew. The same model at xhigh gets twenty-nine of thirty for about twenty cents more, so spend the effort instead.
27 gpt-5.6-luna · xhigh ChatGPT (Codex) 23/30 · 77% ~5 min ~$0.12
Eleven cents a part, and twenty-three of thirty right where the same model at low effort managed five. Effort is what luna was missing. It still slips when a part has to keep working at other sizes: two bit blocks stopped building once the bit got bigger, and a wall clip put the screw through the only place the bundle had to sit. One knob came out with a round bore that spins on the shaft, and once it filed its guess at the unmeasured dimension without marking it as a guess.
28 gpt-5.6-terra · high ChatGPT (Codex) 23/30 · 77% ~3 min ~$0.79
Twenty-three of thirty right, and the wall clip is where it comes apart: four of five left no usable path for the screw and its driver. It also left one cable clip with no hole in its mounting tab, one knob the stem would not enter, and one bit block that stopped building once the bit grew. The curved rest and the missing measurement it got right every time. The same model at xhigh gets twenty-nine of thirty for about twenty cents more a part, which is the whole difference between these two rows.
29 gpt-5.6-luna · high ChatGPT (Codex) 22/30 · 73% ~4 min ~$0.10
Ten cents a part, and twenty-two of thirty right. The job it never got is the bit block: every one built at the size stated and then stopped tracking once the bit grew, losing its pockets or its top chamfer. Elsewhere it slipped once each, a wall clip with no way in for the screw, a rest the pole would not sit in at another size, and a knob with a round bore that spins on the shaft. The same model at xhigh is the same machine for two cents more.
30 gpt-5.6-luna · medium ChatGPT (Codex) 11/18 · 61% ~2 min ~$0.079
Eleven of eighteen right at about two minutes and a dime a part. All three knobs came out too tight for the stem, two wall clips stopped holding the bundle once its size changed, one bit block lost its chamfers when the bit grew, and once it wrote its guess at the unmeasured dimension down as though it had measured it. The same model at xhigh costs pennies more and misses less; run that.
31 gpt-5.6-terra · low ChatGPT (Codex) 10/18 · 56% ~2 min ~$0.74
Everything it made built, and about half were worth printing. The pattern is a part that works at the size you stated and nowhere else: all three of its wall clips stopped fitting when the cable bundle changed, and one pole rest came out flat where the job needed a curve. It never once went back to measure what it had made. The same model at xhigh is a different machine for about a quarter more a part.
32 gpt-5.6-luna · low ChatGPT (Codex) 5/18 · 28% ~3 min ~$0.077
Cheap, fast, and right five times out of eighteen. It wrote the pole's size straight into the file and still built a rest the pole would not drop into, at that size or any other. Elsewhere it left a 0.3mm wall no printer will lay down, and once wrote its guess at the unmeasured dimension down as though it had measured it. The same model at xhigh is a different machine; run that instead.
33 claude-haiku-4-5 · high Claude 2/12 · 17% ~8 min ~$0.46
The weakest row here, and the extra effort did not help. Two parts of twelve came out right. It also wrote its guess at the unmeasured dimension down as though it had measured it, which is the mistake nobody catches until the print is wrong six months later.
34 claude-haiku-4-5 · low Claude 4/33 · 12% ~7 min ~$0.46
Four parts of thirty-three came out right, the bottom of the board. The cable clip, the simplest job here, is the only job it got right more than once, and even the fully spelled-out bit block never once got its chamfers. All three design jobs failed on every attempt: poles that would not drop into the rest, knobs too tight for the stem, wall clips with no way in for the screw and its driver. Twice it wrote its guess at the unmeasured dimension down as though it had measured it. At around forty-six cents a part it costs six times what the best row on the board costs.
Quality against speed.
Every dot is one model at one effort setting: how often its parts printed right, against how long it took per part. The best picks sit high and to the left, right most of the time without the wait.
Six jobs, graded on geometry.
$/part is what the same tokens would cost at API list prices; on a subscription it comes out of your plan.
Early days: 912 graded parts across 6 jobs so far. Each bar averages every attempt on file, and the ticks are the attempts themselves.
Grading is a fixed rubric measured on the part's actual geometry, so the only randomness is the model's. Raw results, full transcripts, and the grading code are on GitHub.