LUNA-TECH.CO← back to luna-tech.co
FIELD NOTE · 005 / PS·ALI AI BAKEOFFLUNA TECH LLC · SEATTLE, WAGENERATED 2026.06.18

Field Note · 005 · Bench

PS/ALI AI Bakeoff: analysis & full aggregate.

E911-style address-validation benchmark · Cohort 1 (local) & Cohort 2 (cloud/API) · report generated 2026-06-18

Each model is fed civic address records (Seattle streets like Madison, Pike, MLK, Stewart, Aurora…; ZIP 98122 etc.) and asked to propose correction probes (expand MadMadison, fix a street suffix STAVE, etc.). Each probe is checked against authoritative reference data, so the harness can tell a validated correction from a guess. This is an E911-style address-validation task (PS/ALI = Private Switch / Automatic Location Identification, the feature that feeds per-extension location from a private switch/PBX into the 911 ALI database of record), which is why “don’t guess” behavior is scored as heavily as raw accuracy. A confidently wrong dispatch address is a safety failure, not just an error.

Summary chart — accuracy vs. calibration vs. contract failures.

ideal: accurate + well-calibrated0.00.20.40.60.80.10.20.30.40.50.60.7Valid probe rate → (higher = more validated corrections)ECE: calibration error (lower / up = better)12345678910RANKED BY VALID RATEfill = JSON-fail % · ring = cohort (1 local / 2 cloud)1gpt-oss-20b (free)VR .77 · ECE .25 · JSON-fail 30%2nemotron-30bVR .74 · ECE .17 · JSON-fail 3%3gemini-flashVR .68 · ECE .34 · JSON-fail 50%4qwen3-coder:30bVR .63 · ECE .22 · JSON-fail 0%5llama-3.1-8bVR .63 · ECE .22 · JSON-fail 3%6llama-3.3-70bVR .61 · ECE .31 · JSON-fail 2%7gemma-4-e4bVR .54 · ECE .38 · JSON-fail 0%8nemotron-9b (free)VR .42 · ECE .38 · JSON-fail 0%9gemini-flash-liteVR .23 · ECE .37 · JSON-fail 0%10llama3.2:3bVR .17 · ECE .66 · JSON-fail 0%

Single-pass runs. X = valid probe rate (right is better) · Y = ECE / calibration error (up is better) · fill colour = share of responses lost to invalid JSON (green→red) · ring = cohort. Points are numbered and ranked in the side panel; coincident markers (e.g. #4 qwen and #5 llama-3.1-8b, nearly identical) are nudged apart with a thin tie-line to their true position. Read the story directly: #2 nemotron-30b and #4 qwensit top-right in green (the win); #1 gpt-oss and #3 gemini-flash are top-right but red: accurate yet emitting a third to a half of responses as malformed JSON the validator discards.

Findings.

Cohort 1: local models (n=10; firm)

  • qwen3-coder:30b, the clear winner. ~63% valid, best ECE (0.22), zero contract violations on single-pass/CF, fast (~6.7s). But it’s aggressive (~28% authoritative-validation-failure + ~4.5% overreach), and its contract violations spike to 37.6 in stability mode.
  • gemma-4-e4b: competent but ~4× slower (~27s) for less accuracy and worse calibration (ECE 0.38); documented JSON-breakage history in .archive/.
  • llama3.2:3b, the floor: ~17% valid, uncalibrated (ECE 0.66), punts on ~69% of records. It is the most consistent model in either cohort, but only because it reliably does nothing useful. Stability ≠ quality.

Cohort 2: cloud / API (mostly n=3; directional)

  • nemotron-3-nano-30b: best cloud model that emits valid output: top validity (~74%), best ECE (0.17), strong reasons (0.51). Cost: ~10s latency and the lowest stability in the cohort (reasoning model).
  • llama-3.1-8b, the workhorse; only full 10/10/5 run, solid + fast + cheap, but terse reasons (0.19) and a contract-violation spike in stability (→98.6).
  • gemini-2.5-flash, effectively broken here: ~48–60% invalid JSON, ~22 retries/run, despite good underlying accuracy. Flash-specific. Flash-Lite (~1% fails) is clean but very conservative.
  • gpt-oss-20b (free): highest raw validity (~77%) but ~30% JSON failures. nemotron-9b (free) is clean but low-accuracy and slow (free-tier queueing, not the model).

Cross-cutting themes.

Full aggregate tables.

Scalars averaged per target per mode, read from each run’s aggregate JSON block (reproduces Cohort 1’s official report exactly). Cell colour: green = better, red = worse, per each column’s ↑/↓.

Single-pass

Provider:ModelRunsRecsValidatedOverreachProbe Eff ↑Valid Rate ↑ECE ↓Contract Viol ↓Lat ms ↓RetrReason Q ↑
1 lmstudio:google/gemma-4-e4b10402800.5360.5360.37902692000.41
1 ollama:llama3.2:3b1040910.1740.1740.6621.5241200.23
1 ollama:qwen3-coder:30b1040433.90.630.630.2233.9670400.49
2 google:gemini-2.5-flash-lite3202200.2280.2280.369019390.30.55
2 google:gemini-2.5-flash32010.300.6850.6850.337101071422.30.36
2 nvidia:meta/llama-3.1-8b-instruct1040160.44.50.6270.6270.2195.955060.20.19
2 nvidia:meta/llama-3.3-70b-instruct3408100.610.610.3140.7219963.30.38
2 nvidia:nvidia/nemotron-3-nano-omni-30b-a3b-reasoning34029.700.7410.7410.16721050930.5
2 openrouter:nvidia/nemotron-nano-9b-v2:free3209.300.4190.4190.380371831.70.44
2 openrouter:openai/gpt-oss-20b:free320210.70.7690.7690.2526.71647480.44

Counterfactual

Provider:ModelRunsRecsValidatedOverreachProbe Eff ↑Valid Rate ↑ECE ↓Contract Viol ↓Lat ms ↓RetrReason Q ↑Sens ↓
1 lmstudio:google/gemma-4-e4b104027.700.5290.5290.3490.52665500.410.725
1 ollama:llama3.2:3b10407.91.20.1690.1690.6761.2220500.20.525
1 ollama:qwen3-coder:30b104042.84.30.6270.6270.2334.3656300.50.75
2 google:gemini-2.5-flash-lite3201800.2070.2070.3950.320411.30.540.75
2 google:gemini-2.5-flash3201100.7040.7040.4239.3932122.30.390.667
2 nvidia:meta/llama-3.1-8b-instruct1040158.86.10.590.590.2267.958750.10.20.55
2 nvidia:meta/llama-3.3-70b-instruct3406700.5120.5120.2280.7284055.70.370.417
2 nvidia:nvidia/nemotron-3-nano-omni-30b-a3b-reasoning34029.30.30.6470.6470.245383762.30.510.583
2 openrouter:nvidia/nemotron-nano-9b-v2:free320900.3940.3940.3710433872.70.450.583
2 openrouter:openai/gpt-oss-20b:free32021.30.70.7280.7280.1486141729.30.440.833

Stability

Provider:ModelRunsRecsValidatedOverreachProbe Eff ↑Valid Rate ↑ECE ↓Contract Viol ↓Lat ms ↓RetrReason Q ↑Dec Cons ↑Cand Cons ↑Conf Var ↓
1 lmstudio:google/gemma-4-e4b537624300.4540.4540.4344.2293385.20.4185.1%87.2%0.061
1 ollama:llama3.2:3b537676.66.60.1670.1670.6846.8225100.2197.1%99.2%0.005
1 ollama:qwen3-coder:30b5376391.437.60.610.610.2437.6658900.4994.0%94.3%0.031
2 nvidia:meta/llama-3.1-8b-instruct53761420.271.80.5890.5890.22498.646680.20.1985.1%91.4%0.04
2 nvidia:meta/llama-3.3-70b-instruct2376619.500.5120.5120.2151.522385240.3892.1%94.4%0.013
2 nvidia:nvidia/nemotron-3-nano-omni-30b-a3b-reasoning237625350.6440.6440.22916.58828130.5274.7%82.4%0.09

Decision-outcome mix (single-pass + counterfactual, % of records)

Provider:Modelacceptaccept+normsuggest_reviewneeds_reviewvalid_failcontract_failoverreach_fail
1 lmstudio:google/gemma-4-e4b22.5%23.8%5.8%42.0%5.9%0.1%0.0%
1 ollama:llama3.2:3b20.4%7.9%0.1%68.8%0.0%0.1%2.8%
1 ollama:qwen3-coder:30b22.5%32.4%4.5%8.6%27.5%0.0%4.5%
2 google:gemini-2.5-flash-lite15.0%29.2%4.2%50.8%0.0%0.8%0.0%
2 google:gemini-2.5-flash5.0%20.8%1.7%21.7%2.5%48.3%0.0%
2 nvidia:meta/llama-3.1-8b-instruct22.1%3.4%29.5%37.9%0.0%3.4%3.8%
2 nvidia:meta/llama-3.3-70b-instruct22.5%12.9%31.2%31.7%0.0%1.7%0.0%
2 nvidia:nvidia/nemotron-3-nano-omni-30b-a3b-reasoning22.1%36.7%12.1%25.4%0.0%3.3%0.4%
2 openrouter:nvidia/nemotron-nano-9b-v2:free15.0%18.3%10.0%56.7%0.0%0.0%0.0%
2 openrouter:openai/gpt-oss-20b:free12.5%34.2%3.3%18.3%0.0%28.3%3.3%

Coverage & caveats.

Source.

Harness, prompts, and judge live at github.com/lunatechsea/PSALI-Bakeoff. Run JSON is kept locally; the tables above are the aggregate scalars each run emits, which regenerate from the same harness against any model you point it at.

Validated / Records are per-run totals on different record sets (Cohort 1 + NVIDIA = 40-record set; Gemini/OpenRouter = 20; stability = 376). Use the normalized Probe Eff / Valid Rate for fair cross-model comparison, not raw Validated.