September 29, 2026 · 5:35 · Evidence through September 24
14 local AI setups on a real app
Signal stood out on one Swift parser task. Here are the outcomes, recorded settings, independent checks, and files behind the video.
Watch on YouTube: I Tested 14 Local AI Setups on a Real App — Signal Stood Out
What the test showed
7 / 14
local setups passed the reported task criteria
Signal Q6: 11m21s, 91/91 tests, Debug build passed. The frozen patch passed independent evaluation without evaluator repairs. This was the fastest passing local observation.
The task added U+2011 non-breaking hyphen support to verse ranges, normalized accepted separators, and preserved visible Markdown text and decoded URL references. Four authored regression tests accompanied two production-line changes.
Self-written tests were insufficient: older Qwopus 3.6 Q4 passed its own tests using the wrong Unicode character. Fixed independent acceptance checks caught the mistake.
Compare the recorded runs
Bar length shows recorded duration, not quality or a controlled speed ranking. A fast failed run did not complete the task. Captured spans and termination times are labeled separately; missing times have no bar. Select a setup for its settings and caveats.
Showing 14 of 14 setups · 7 reported passes
Selected setup
Signal 3.8 27B · Q6
AP-Q6_K GGUF · medium · Pass
- Runtime
- Standalone llama.cpp 2.37.0 / Hermes
- Recorded duration
- 11m 21s · recorded run
- Settings
- 118,016 context · q8_0 K/V · native MTP max 4 · backend draft sampling off · BF16 vision projector loaded · temperature .6 / top-p .95 / top-k 20 / min-p .05 / presence penalty 1.0 · same harness/runtime settings as Signal IQ3_S
- Evaluation
- Tests: 91 · failures: 0 · fixed acceptance failures: 0 · build: Pass
What to keep in mind
Fastest passing local observation: 11m21s, 28.7% less time than Signal IQ3_S and 42.5% less than GSQ-RCO. 24 tool calls versus 36 and 22 respectively. Simpler two-line production patch; four authored tests with all-dash full Markdown links and direct raw-variant URL round-trips. Recovered from four URL-encoding assertion failures; corrected own 89-test suite passed, then independent 91/91 and build passed without repair. File search instead of graph, production-before-test ordering, /tmp diagnostics and piped exit-code handling remain workflow limitations. Unrelated TypeScript check lacked tsc and was disclosed. One run per configuration; cache, machine state and work trajectory can affect timing.
Comparison scope and caveatsOne real Swift parser task, with different models, settings, dates and runtimes. Reported pass status and timing coverage are separate. The complete evidence entries and downloadable results remain below.
All 14 local setups
Each entry is a complete setup, including runtime and reasoning tier. Open an entry for its settings and evaluation. Unknown values remain unknown.
Qwen3.8 27B
Q6_K GGUF · medium · 2026-08-24
Pass52m 40sSettings and evidence
Qwen3.8 27B
Q6_K GGUF · medium · 2026-08-24
Correct two-file patch. Repeated planning and test launches delayed completion; first write took nearly 35 minutes.
Runtime: LM Studio / llama.cpp
Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens
Independent combined suite: 92/92 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.
Grug v1.1 27B
Q6_K GGUF · medium · 2026-08-24
Fail5m 22sSettings and evidence
Grug v1.1 27B
Q6_K GGUF · medium · 2026-08-24
Fast but missed U+2011; inserted U+200B removal and repeated ASCII hyphen. Wrote tests but did not run them.
Runtime: LM Studio / llama.cpp
Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens
Independent combined suite: 78/90 passed. Fixed acceptance assertion failures: 8. Debug build: Pass.
Kiwen1.1 27B
Q6_K GGUF · medium · 2026-08-25
Pass88m 43sSettings and evidence
Kiwen1.1 27B
Q6_K GGUF · medium · 2026-08-25
Eventually correct after a long regex repair loop. 68 tool calls and 88.7 minutes; also expanded support to U+2010 and wrote scratch files outside its worktree.
Runtime: LM Studio / llama.cpp
Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens
Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.
Grug v1.1 27B
Q6_K GGUF · xhigh · 2026-08-25
Fail4m 23sSettings and evidence
Grug v1.1 27B
Q6_K GGUF · xhigh · 2026-08-25
Verified native xhigh did not repair correctness. Recognized U+2011 but failed normalization; did not run its useful regression test.
Runtime: LM Studio / llama.cpp
Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens
Independent combined suite: 69/88 passed. Fixed acceptance assertion failures: 12. Debug build: Pass.
Qwen3.8 27B
4-bit MLX · medium · 2026-08-26
No patch69m 15sSettings and evidence
Qwen3.8 27B
4-bit MLX · medium · 2026-08-26
Diagnosed the intended fix but never applied it. Two Metal out-of-memory retries, unsupported structured requests, and long non-tool responses. Only wrote a scratch harness.
Runtime: mlx_vlm.server
Settings: 262,144 loaded context · MTP max 3 · KV 8-bit / group 64, starts at token 5,000
Independent combined suite: 71/87 passed. Fixed acceptance assertion failures: 16. Debug build: Pass.
Qwopus3.6 35B A3B
oQ4 MLX · medium · 2026-08-26
Fail33m 42sSettings and evidence
Qwopus3.6 35B A3B
oQ4 MLX · medium · 2026-08-26
Its own 92 tests passed, but production code and tests shared the same wrong character: U+200C instead of U+2011. Independent checks caught it. Two Xcode diagnostic timeouts also increased wall time.
Runtime: mlx_vlm.server
Settings: 118,016 Hermes context / 262,144 model context · extracted native MTP max 3 · KV 8-bit / group 64
Independent combined suite: 86/94 passed. Fixed acceptance assertion failures: 8. Debug build: Pass.
Muse Glimmer 30B
6-bit MLX · Not verified · 2026-08-28
No patch47m 27s captured spanSettings and evidence
Muse Glimmer 30B
6-bit MLX · Not verified · 2026-08-28
Retained usage and session evidence show 16 model calls without a usable patch or final response. The later status audit confirmed a clean benchmark worktree. DFlash was not exercised. Full independent Xcode results were not recovered.
Runtime: MLX / Hermes
Settings: Base MLX path; no DFlash drafter. Exact effective context and rendered reasoning tier not verified in retained summary.
Independent combined suite: Not recovered. Fixed acceptance assertion failures: Not recovered. Debug build: Not recorded.
Qwen3.8 + DFlash2 · 118K
6-bit MLX · medium · 2026-08-30
Pass27m 15sSettings and evidence
Qwen3.8 + DFlash2 · 118K
6-bit MLX · medium · 2026-08-30
Passed with 24 tool calls, no exact repeated calls, and no recorded context compaction. 48.3% less wall time than the historical Q6 control. About 12.88 GiB observed swap and seven memory-guard cache sheds remained. DFlash draft acceptance: 56.50%. This end-to-end speed difference is not an isolated on/off measurement of speculation.
Runtime: mlx-dspark / DFlash2
Settings: 118,016 context · DreamFoundries/Qwen3.8-27B-6bit + incoai/Qwen3.8-27B-DFlash2 · medium
Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.
Qwen3.8 GSQ-RCO 27B
IQ3_S GGUF · medium · 2026-09-13
Pass19m 44sSettings and evidence
Qwen3.8 GSQ-RCO 27B
IQ3_S GGUF · medium · 2026-09-13
Correct two-line parser change and four authored tests. Recovered from an obsolete Xcode path and a faulty URL assertion; final independent suite passed. Test coverage has minor gaps described in the run report. Previously fastest passing local configuration; Signal later completed this task in 15m56s.
Runtime: LM Studio / llama.cpp 2.37.0
Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens
Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.
Signal 3.8 27B
AP-IQ3_S GGUF · medium · 2026-09-16
Pass15m 56sSettings and evidence
Signal 3.8 27B
AP-IQ3_S GGUF · medium · 2026-09-16
Earlier passing local observation: 15m56s, 19.3% less time and 47.7% fewer output tokens than GSQ-RCO. More tool calls (36 vs 22), so speed did not mean fewer actions. Four authored tests; recovered from an invalid Xcode selector and two URL-encoding assertion failures. Explicit Unicode inspection and exact Markdown links were strengths. Used file search instead of graph; wrote a /tmp scratch file despite the scope constraint; edited production before tests. Independent 91/91 and build passed without repair. Different endpoint wrapper, q8 KV and MTP max 4 mean this is a complete-configuration comparison, not an isolated fine-tune effect.
Runtime: Standalone llama.cpp 2.37.0 / Hermes
Settings: 118,016 context · q8_0 K/V · native MTP max 4 · backend draft sampling off · vision projector loaded · temperature .6 / top-p .95 / top-k 20 / min-p .05 / presence penalty 1.0 · Xcode 27 beta 6
Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.
Signal 3.8 27B · Q6
AP-Q6_K GGUF · medium · 2026-09-17
Pass11m 21sSettings and evidence
Signal 3.8 27B · Q6
AP-Q6_K GGUF · medium · 2026-09-17
Fastest passing local observation: 11m21s, 28.7% less time than Signal IQ3_S and 42.5% less than GSQ-RCO. 24 tool calls versus 36 and 22 respectively. Simpler two-line production patch; four authored tests with all-dash full Markdown links and direct raw-variant URL round-trips. Recovered from four URL-encoding assertion failures; corrected own 89-test suite passed, then independent 91/91 and build passed without repair. File search instead of graph, production-before-test ordering, /tmp diagnostics and piped exit-code handling remain workflow limitations. Unrelated TypeScript check lacked tsc and was disclosed. One run per configuration; cache, machine state and work trajectory can affect timing.
Runtime: Standalone llama.cpp 2.37.0 / Hermes
Settings: 118,016 context · q8_0 K/V · native MTP max 4 · backend draft sampling off · BF16 vision projector loaded · temperature .6 / top-p .95 / top-k 20 / min-p .05 / presence penalty 1.0 · same harness/runtime settings as Signal IQ3_S
Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.
Qwopus3.8 27B Flash V2 · Q6
MLX affine 6bit · group 64 · medium · 2026-09-22
Pass17m 03sSettings and evidence
Qwopus3.8 27B Flash V2 · Q6
MLX affine 6bit · group 64 · medium · 2026-09-22
Passed in 17m03s: 91 independent tests and Debug build, without external patch repair. Two production lines and four authored tests. 16 tools and 5,619 output tokens, fewer than Signal Q6, but 50.2% more elapsed time. Used file search instead of graph and production-before-test ordering; recovered unaided from stale Xcode path and test-without-building before a build. Authored Markdown and URL checks are narrower than Signal Q6; fixed independent acceptance tests passed. Runtime, KV representation, no MTP and released Xcode differ from Signal. Single observation, not an isolated model-quality comparison. Vision passed a basic image smoke test only.
Runtime: mlx-vlm 0.6.15 / Hermes
Settings: 118,016 context · MLX 8-bit KV group 64 from token 0 · prompt cache on and observed · no MTP drafter · vision included and smoke-tested · temperature .6 / top-p .95 / top-k 20 / min-p .05 / presence penalty 1.0 (64-token window) · verified sampler correction · Xcode 27.0 released
Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.
ThinkingCap Qwen3.8 27B
Q6_K GGUF · medium · 2026-09-24
Agent failure60m 15s to terminationSettings and evidence
ThinkingCap Qwen3.8 27B
Q6_K GGUF · medium · 2026-09-24
Incomplete: stopped by the controller after 60m 15s with no final agent response. Independent evaluation of the unmodified patch passed both fixed acceptance methods (eight dash/spacing cases) and the Debug build. The combined suite passed 90 of 91 tests; the sole failure was a model-authored Markdown test incorrectly expecting colon percent-encoding (%3A). The one-hour cutoff was chosen around minute 48, not preregistered or uniformly applied to earlier runs. Elapsed time at termination is not a successful completion time.
Runtime: LM Studio bundled llama.cpp 2.37.0, commit 8172e65; standalone server
Settings: Hermes 3aee290899e478c5fdfb6a241ef62758a49829b3; 118,016 context; q8_0 K/V; native MTP max 4; backend draft sampling disabled; temperature 1, top-p 0.95, top-k 20, min-p 0, presence penalty 0. Released Xcode 27; iPhone 17 / iOS 26.5. Same historical baseline, prompt and fixed evaluator. Vision projector loaded and image smoke test passed; coding benchmark was text-only
Independent combined suite: 90/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.
Ornith 1.5 35B-A3B
Not recovered · Not recovered · Date not recovered
Reported failNot recoveredSettings and evidence
Ornith 1.5 35B-A3B
Not recovered · Not recovered · Date not recovered
The user-provided historical comparison lists Ornith 1.5 35B-A3B as failed. Detailed benchmark logs, timing, quantization, context, and test counts were not recovered. This is a reported result, not an independently revalidated measurement. It is not conflated with a separately found Ornith 9B download.
Runtime: Not recovered
Settings: Configuration named in the supplied screenshot; detailed runtime settings unavailable
Independent combined suite: Not recovered. Fixed acceptance assertion failures: Not recovered. Debug build: Not recorded.
The 14 include the historical user-reported Ornith result, explicitly unverified. Gemini (14m07s) and Luna/max (6m50s) passed as hosted references outside this local count. The separate DFlash Sharp diagnostic failed the agent task despite external test passes. Bonsai was cancelled, incomplete, and unscored.
Signal Q6 recorded settings
- Weights
- Signal 3.8 27B AP-Q6_K GGUF
- Runtime
- Standalone llama.cpp 2.37.0 / Hermes
- Context capacity
- 118,016 tokens
- KV precision
- q8_0 K/V
- Native MTP
- Max 4; backend draft sampling off
- Sampler
- Temperature .6; top-p .95; top-k 20; min-p .05; presence penalty 1.0
These are observed run settings, not a tested installation recipe. Weight quantization, KV precision, context capacity, prompt-prefix reuse, and MTP are separate controls. No isolated MTP speedup was measured.
Take the findings with you
Curated data and settings, with checksums. No private app source, patches, raw logs, credentials, or model weights are included.
How far these results travel
One bounded parser task on Jason’s MacBook Pro M5 Pro with 48 GB memory. This is not a general model ranking. The combined test counts include baseline tests, model-authored tests, and two fixed acceptance methods; 91/91 does not mean 91 independent coding tasks.
Agent wall time excludes model loading, warm-up, baseline preparation, and independent evaluation. Runtime, harness, reasoning tier, cache state, Xcode version, and work trajectory differ. A single observation per setup does not isolate quantization, fine-tuning, or context size as a cause.
ThinkingCap was stopped after 60m15s without a final answer. Both fixed acceptance methods and the build passed, but one authored test failed. Its termination time is not a completion time; the cutoff was chosen during the run, not preregistered.
Supplementary DFlash 65K and 118K runs both passed. The 65K run reported 95% context occupancy; the 118K run recorded zero compactions. Both showed memory pressure and cache trimming. This supports a capacity-headroom observation, not a correctness advantage.
Evidence identity
The curated export derives from the immutable September 24 film ledger. All 66 indexed source-file hashes matched the retained archive when prepared. Raw receipts remain private.
Ledger SHA-256: b1fc4463de731a470b42f189f9b51863c036ac38a34f43c57d954e920c2dfded
Fixed evaluator SHA-256: f8c54b591bed62ee5fd7eec5e81e7c9e88b6481511b726369f9f31e768fc7f64