AI Efficiency Toolbox
Local AI · findings first

What should you use—and why?

Choose your hardware and work. Compare useful results, failure modes and exact setups before you download.

Reviewed 2026-10-02
4 artifacts · selected coverage
Persistent evidence database

Optional browser hints, read on this device only. Your current and saved choices stay unchanged until you confirm. You can always select manually below.

Manual choices remain available. No model inference runs.

Select hardware and memory to estimate fit. Below are task-relevant evidence leads. No complete hardware/quant/runtime-version match in this collection. Results below are qualified leads, not hardware-validated winners.

Best-supported leads

3 supported leads · 2 community reports for this task
Browse 11 new community investigations

Evidence at a glance

Community outcome counts

Related-model reports across different setups. Each bar shows outcome shares, not total-report magnitude; N is the denominator. Click to inspect sources.

Signal 3.8 · AP-Q6_KN=0

No eligible community reports. Firsthand evidence is separate.

Terse-Coder 27B Q6_K · native MTPN=0

No eligible community reports. Firsthand evidence is separate.

Qwen3.8 27B · UD-Q5_K_MN=2

Weight-file memory

Set memory to compare weights with your budget.

Signal 3.8 · AP-Q6_K20.9 GiB

Memory fit unknown

Terse-Coder 27B Q6_K · native MTP20.9 GiB

Memory fit unknown

Qwen3.8 27B · UD-Q5_K_M18.4 GiB

Memory fit unknown

File sizes are pinned metadata, not measured RAM. Planning reserve: 6 GiB for Apple/CPU, 2 GiB for GPU VRAM. KV cache, context, projector, workspace and offloading are not modeled. A weight fit does not establish practical capability.

Lead 1 · qualified evidenceAP-Q6_K

Signal 3.8 · AP-Q6_K

Choose hardware to assess fit

Strength
Recovered from four URL-encoding assertions; completed a small production patch.
Likely struggle
File-search workflow, temporary diagnostics and test ordering remain limitations.
Memory
20.9 GiB weights · Memory fit unknown

0 independent community reports: 0 success · 0 partial · 0 failure · 0 unclear. Plus 1 separate firsthand observation.

Settings, artifact identity & publisher claim

Recorded: temperature .6, top-p .95, top-k 20, min-p .05, presence penalty 1; q8_0 K/V; native MTP max 4 and backend draft sampling off. Recorded 118,016 context is not a safe default for every device.

Settings source · apache-2.0

AP-Q6_K/Signal-3.8-27B-AP-Q6_K.gguf
Revision: ae272aa3e533854feebcfa1ded0f1dca9afedcef
SHA-256: fe4760215d62b697686a7e608a9705b629003103228228b37e104da22e2385c0

Pinned download metadata checked October 1. Historical reports do not verify their bytes against this SHA-256; matching names/quantizations are not proof of identical files.

Publisher/author claim: Author emphasizes faster useful answers and provides sampler/cache guidance. No numerical reduction claim is entered here. Source

Claim is attributed, not independently confirmed; it does not determine ordering. No comparable vendor/community thinking-reduction percentage is available.

27B total parameters · dense. Full stored weights determine memory.

AI Efficiency Toolbox evidence review · 2026-10-01 · Qualified lead. Jason’s strongest observed local coding result in the published comparison. One recorded run, no independent community coverage in this selection; hardware applicability unverified.

  • Initial curated review, October 1, 2026. No representative survey or new benchmark.
Copy an agent setup handoff

Includes your selected hardware/task and the exact artifact. Your agent must check resources and its own permissions; without computer tools it can guide manual setup.

Jason’s separate firsthand observation

Success · Jason

Completed the real app task; independent 91/91 tests and build passed.

2026-09-17 · Local Mac; chip/memory not specified in public export · memory unknown · AP-Q6_K · llama.cpp 2.37.0

  • Worked: Recovered from four URL-encoding assertions; completed a small production patch.
  • Struggled: File-search workflow, temporary diagnostics and test ordering remain limitations.
  • Correction: Corrected its own authored tests before the separate independent check.

Hermes, one observation per setup; text-only coding. 14 local setups / 7 passing overall; no hardware-wide success rate.

Reported measurement: Wall time 681.2 seconds. Single Signal AP-Q6_K observation, 2026-09-17. Not comparable to unrelated community tasks.

Execution: local-only · Planner: none reported. Hermes tool-calling app task; no cloud planning step

Evaluator: Independent 91/91 tests and Debug build without repair. Delivery coverage: requirements partial; implementation checked; independent tests checked; review not-checked; release checks partial. This is not release certification.

Elapsed: 681.1753717500251 seconds · Output tokens: 7798 · Cost: unknown (no API-equivalent estimate) · Human corrections: unknown

Read source

Source checked 2026-10-01

Lead 2 · qualified evidenceQ6_K

Terse-Coder 27B Q6_K · native MTP

Choose hardware to assess fit

Strength
Verse-range separator/spacing canonicalization, original visible Markdown preservation, URL round trips and ambiguity protection passed fixed focused evaluation. Tests-first workflow and unaided recovery from stale Xcode path observed.
Likely struggle
No acceptance-test/build failure in this completed run. Instruction-following caveats: obsolete Xcode path initially, output pipelines without pipefail and default DerivedData outside worktree.
Memory
20.9 GiB weights · Memory fit unknown

0 independent community reports: 0 success · 0 partial · 0 failure · 0 unclear. Plus 1 separate firsthand observation.

Settings, artifact identity & publisher claim

Medium reasoning; context 118016; q8_0 K/V; native draft-mtp max 4, backend draft sampling disabled; temperature 0.6, top-p 0.95, top-k 20, min-p 0.05, presence penalty 1.0; response cap 8192; parallel 1; no vision projector. Hermes 3aee290899e478c5fdfb6a241ef62758a49829b3. Author template Shockem/froggeric-terse-coder at c557b014b6675a87683b45d0800a68842311ffe2 (rendered anti-rumination/medium-effort use verified). Qualify MTP/tool parsing on the exact runner before relying on these settings.

Settings source · Apache-2.0

Qwen3.8-27b-Terse-Coder.Q6_K.gguf
Revision: 7602eadf6a90853492e4e8a377c2a2f531bf21e2
SHA-256: 3164a8d8cc7f5c669b7a6de57ce9e5db126d73b8d39a0f424a2b9c649f25df07

Actual local model bytes rehashed and size checked against pinned download: SHA256 3164a8d8cc7f5c669b7a6de57ce9e5db126d73b8d39a0f424a2b9c649f25df07. Author template separately pinned; runner path/alias and rendered template proof verified. This exact first-party run does not establish equivalence with other quants/runtimes.

Publisher/author claim: Shockem identifies this pinned artifact as Qwen3.8-27b-Terse-Coder Q6_K and licenses the repository Apache-2.0. No publisher speed/quality claim is used for ordering. Source

Claim is attributed, not independently confirmed; it does not determine ordering. No comparable vendor/community thinking-reduction percentage is available.

27B total parameters · dense. Full stored weights determine memory.

AI Efficiency Toolbox · 2026-10-02 · Personally tested lead; community coverage absent. One independently checked first-party FaithVoice parser run passed on Apple M5 Pro 48GiB. Exact bytes/settings are documented; no cross-variant quality transfer, isolated MTP speedup or release qualification established. Keep this separate from the historical 14-setup study and community observations.

  • 2026-10-02: add verified latest native-MTP run; earlier no-MTP pass retained as context, not another independent community report.
Copy an agent setup handoff

Includes your selected hardware/task and the exact artifact. Your agent must check resources and its own permissions; without computer tools it can guide manual setup.

Jason’s separate firsthand observation

Success · Jason

PASS: 92 independent focused parser tests, zero failures; independent Debug build passed. Agent elapsed 11m23.131s, 7427 output tokens.

2026-10-02 · Apple M5 Pro / Mac17,8; same-Mac local inference, hardware corroborated post-run · 48 GiB · Q6_K · llama.cpp 2.37.0

  • Worked: Verse-range separator/spacing canonicalization, original visible Markdown preservation, URL round trips and ambiguity protection passed fixed focused evaluation. Tests-first workflow and unaided recovery from stale Xcode path observed.
  • Struggled: No acceptance-test/build failure in this completed run. Instruction-following caveats: obsolete Xcode path initially, output pipelines without pipefail and default DerivedData outside worktree.
  • Correction: No evaluator-guided repair; original patch preserved before fixed evaluation. Earlier no-MTP run also passed (91 tests/build, 783.378s, 5972 output tokens, 22 tools). Single-run trajectories and changed historical Signal toolchain/proxy/search/projector prevent isolated speedup or quality claims. No merge/release/device qualification; reasoning count and operating cost unknown.

Medium reasoning; context 118016; q8_0 K/V; native draft-mtp max 4, backend draft sampling disabled; temperature 0.6, top-p 0.95, top-k 20, min-p 0.05, presence penalty 1.0; response cap 8192; parallel 1; no vision projector. Hermes 3aee290899e478c5fdfb6a241ef62758a49829b3. Author template Shockem/froggeric-terse-coder at c557b014b6675a87683b45d0800a68842311ffe2 (rendered anti-rumination/medium-effort use verified). Qualify MTP/tool parsing on the exact runner before relying on these settings. Baseline 2c7aebd67d44e4858f94dce73de5c4f99173fc57; prompt SHA256 a9bd80dad2cf9792ebaacdde38e71a2a0b4c85113f578954c5ced789a9710a9d. MTP smoke verified text, parsed tool call and roundtrip; benchmark drafts accepted 5332/8536 (62.46%), excluding warm-up. Tool calls 19; API calls 19.

Reported measurement: Agent elapsed time 683.1 seconds. One FaithVoice parser task on pinned baseline/prompt; 92 focused tests/build independently checked. Historical 14-setup video data is not replaced.

Execution: local-only · Planner: none reported. Fixed parser task/pinned baseline with Hermes local-model tool workflow; original patch captured before independent acceptance fixture; no evaluator-guided repair.

Evaluator: Independent fixed acceptance fixture and separate focused Xcode tests/Debug build; not model self-asserted success.. Delivery coverage: requirements checked; implementation checked; independent tests checked; review partial; release checks not-checked. This is not release certification.

Elapsed: 683.1314880419998 seconds · Output tokens: 7427 · Cost: unknown (no API-equivalent estimate) · Human corrections: unknown

Read source

Source checked 2026-10-02

Lead 3 · qualified evidenceUD-Q5_K_M

Qwen3.8 27B · UD-Q5_K_M

Choose hardware to assess fit

Strength
UD-Q5_K_M identified a metadata bug and retested.
Likely struggle
Reasoning budget dominated; small unequal sample and LLM judgment limit conclusions.
Memory
18.4 GiB weights · Memory fit unknown

2 independent community reports: 1 success · 1 partial · 0 failure · 0 unclear.

Settings, artifact identity & publisher claim

Publisher documents reasoning_effort / preserve_thinking controls. Begin with a bounded task and verify template/runtime support; medium is a task-specific comparison candidate, not a universal optimum.

Settings source · apache-2.0

Qwen3.8-27B-UD-Q5_K_M.gguf
Revision: 4ca720788d1e01f1bff70c033e0d0028fd02e502
SHA-256: 2de73110cb254cbf09b54b717578dadff12ef1194e7271527e68202f39ba4bfd

Pinned download metadata checked October 1. Historical reports do not verify their bytes against this SHA-256; matching names/quantizations are not proof of identical files.

Publisher/author claim: Publisher claims improved coding and long-horizon agent execution, with flexible thinking control. Source

Claim is attributed, not independently confirmed; it does not determine ordering. No comparable vendor/community thinking-reduction percentage is available.

27B total parameters · dense. Full stored weights determine memory.

AI Efficiency Toolbox evidence review · 2026-10-01 · Qualified lead. One AMD tool-building success and one different-quant Mac overthinking report. Related-family evidence is useful context, not confirmation that this exact artifact works on your hardware.

  • Initial curated review, October 1, 2026. No representative survey or new benchmark.
Copy an agent setup handoff

Includes your selected hardware/task and the exact artifact. Your agent must check resources and its own permissions; without computer tools it can guide manual setup.

Source-backed community findings

Showing 2 of 2 eligible independent reports for this task. Collection window 2026-02-28–2026-10-02. One reported investigation/follow-up group counts once; quoted/reposted runs and irrelevant replies are excluded. Unclear is a separate outcome. This selected denominator is not all users, a representative success rate, or a controlled benchmark.

success · Qwen3.8-27B · FLvdW

Success · FLvdW

Four quantizations produced a working read-only partition checker in one author’s investigation.

2026-09-10 · MI50 32GB + RX 9060 XT 16GB; separate GPU memories · 32 GiB · UD-Q5_K_M · llama.cpp (version unknown)

  • Worked: UD-Q5_K_M identified a metadata bug and retested.
  • Struggled: Long reasoning and truncated tool arguments needed generous budgets.
  • Correction: Raised output budget to 24K and timeout to 900 seconds.

ROCm and custom bash/read/write harness; same-author quant runs count once. Logs offered on request, not independently reproduced.

Execution: local-only · Planner: none reported. Author-described local workflow; complete handoff protocol not supplied

Evaluator: Author self-report; no independent acceptance suite reproduced here. Delivery coverage: requirements partial; implementation partial; independent tests not-checked; review not-checked; release checks not-checked. This is not release certification.

Elapsed: unknown · Output tokens: unknown · Cost: unknown (no API-equivalent estimate) · Human corrections: unknown

Read source

Source checked 2026-10-01

partial · Qwen3.8-27B · TomsnKa

Partial · TomsnKa

Breakout games worked; xhigh used substantially more time with little change in the author’s judged quality.

2026-09-10 · Apple M5 Max; memory not stated · memory unknown · oQ4e-mtp · runtime unknown (version unknown)

  • Worked: Low/medium produced playable games.
  • Struggled: Reasoning budget dominated; small unequal sample and LLM judgment limit conclusions.
  • Correction: Explicitly compare medium to the template’s xhigh default on your task.

3 low / 4 medium / 17 xhigh runs, same prompt; different MLX quant, runtime version absent. One independent report, not 24 users.

Reported measurement: Author-reported wall-time ratio 11 ×. xhigh versus medium on this M5 Max / oQ4e-mtp Breakout task only. Rounded author ratio; no cross-task average.

Execution: local-only · Planner: none reported. Author-described local workflow; complete handoff protocol not supplied

Evaluator: Author self-report; no independent acceptance suite reproduced here. Delivery coverage: requirements partial; implementation partial; independent tests not-checked; review not-checked; release checks not-checked. This is not release certification.

Elapsed: unknown · Output tokens: unknown · Cost: unknown (no API-equivalent estimate) · Human corrections: unknown

Read source

Source checked 2026-10-01

Community field reports

7 setup candidates · 4 contextual observations

Author-reported work across different models, quants and runtimes. These reports do not enter hardware recommendations, fit estimates or outcome charts. Missing artifact identities remain unknown; results do not transfer to another variant. No pooled success rate or independent reproduction is claimed. The existing boutell investigation is enriched once, rather than counted again.

Showing 11 of 11 investigations. Filters apply only to these reports.

Three real tickets on M3 Pro, with browser-error feedback and thinking loopsmouseofcatofschrodi · 2026-04-21 · Setup candidate

MacBook Pro M3 Pro · 36 GiB reported · oMLX · version unknown · harness Pi

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.6-35B-A3B · partial: One ticket completed in one shot using about 7K tokens; two needed browser console errors fed back. Author reports all three solved, but thinking could loop after code was completed.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Three unspecified tickets in an existing real project
  • configured tokens: 126000
  • output limit tokens: unknown
  • kv cache: unknown
  • precision note: 126K as reported, exact token limit unknown
  • temperature: 0.6
  • top p: 0.95
  • top k: 20
  • min p: 0
  • presence penalty: 0
  • repetition penalty: 1
  • preserve thinking: Author tried on, with/without force; loops remained
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.6-35B-A3B · quant MLX 4-bit · artifact unknown · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Author-reported measurements

  • ticket duration minutes: 5–8. Author follow-up gives average range across three tickets.

Limitations and corrections

  • Ticket descriptions and code artifacts not published
  • Initial model typo corrected by author
  • No controlled comparison of harnesses
  • Temperature 1 and preserve_thinking variations did not remove loops for this author

Read mouseofcatofschrodi's source · checked 2026-10-02

M1 Max sales-analysis workflow completed by one Qwen3.5 modelluke_pacman · 2026-02-28 · Setup candidate

M1 Max · 64 GiB reported · llama.cpp server · version build 8179 · harness Author's LangGraph-based agentic application

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.5-35B-A3B · success: Author reports complete analysis, code, visualizations and conclusions using one model; preferred result to previous Nemotron reasoning plus Qwen3-Coder coding setup.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Analyze January 2025 Amazon sales workbook with six sheets, create pandas analysis/visualizations, suggest improvements
  • configured tokens: unknown
  • output limit tokens: unknown
  • kv cache: unknown
  • thinking: Compared enabled and disabled; other sampling parameters not stated
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.5-35B-A3B · quant Q4_K_XL · artifact unknown · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Author-reported measurements

  • generation tokens per second: 27. Author reported.
  • completion minutes thinking off: 15–20. One workflow.
  • completion minutes thinking on: 35–40. One workflow.

Limitations and corrections

  • Quality judged by workflow/app developer, not blinded or independently audited
  • Suggested sales improvement is a recommendation, not demonstrated revenue increase
  • Underlying workbook and exact prompt not available
  • Prior two-model run is same author, not independent corroboration
  • Primary-thread reply identifies Unsloth as quant publisher; historical revision/file/hash still unverified.

Read luke_pacman's source · checked 2026-10-02

Qwen3.5-27B fixes an RL-training bug with better contextgarg-aayush · 2026-03-30 · Setup candidate

RTX 4090 workstation; MacBook used as client via Tailscale · 24 GiB reported · llama.cpp · version unknown · harness OpenCode

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.5-27B · success: Author says RL-training bug was resolved after more precise prompting/context; reports working tool calls and multi-step Python editing/testing. Loose prompts were less effective.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Generate/edit/test Python scripts
  • Debug an RL-training issue
  • configured tokens: 64000
  • output limit tokens: unknown
  • kv cache: unknown
  • precision note: 64K as reported
  • agent skills: true
  • mcp: Context7 documentation retrieval
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.5-27B · quant 4-bit; exact variant not stated in Reddit report · artifact unknown · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Author-reported measurements

  • generation tokens per second: about 40. Author reported.
  • prefill tokens per second: about 2400. Author reported.
  • vram gb used: about 22. Author reported.

Limitations and corrections

  • Inference on NVIDIA, not locally on MacBook
  • No bug reproducer, patch or test log in Reddit evidence
  • Prompting/documentation tools material
  • Same-author blog not another report
  • RTX4090 performs inference; MacBook is a remote client, not Apple inference hardware. Exact 4-bit quant remains unspecified.

Read garg-aayush's source · checked 2026-10-02

Supporting source: Same author's linked setup blog; not another independent report

Devstral gives partial help with unpublished NumPy/Numba RL codeThe_Paradoxy · 2026-03-19 · Contextual observation

Personal RTX 4060 Ti system · 16 GiB reported · Runtime unknown · version unknown · harness Manual chat/code copy-paste; no IDE

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • mistralai/Devstral-Small-2-24B-Instruct-2512 · partial: Only tested model with response author considered partially correct/usable after selecting good fragments.
  • zai-org/GLM-4.7-Flash · failure: Author reports overnight run without useful answer.
  • Qwen/Qwen3.5-27B · failure: Author reports inadequate understanding of this particular novel scientific-code task.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Explain unpublished reinforcement-learning code using NumPy and numba.jit
  • Extend transitive-inference implementation from five elements to seven
  • configured tokens: 20000
  • output limit tokens: unknown
  • kv cache: unknown
  • comparison context note: 20K–48K across models; 20K Devstral
  • cpu offload: Author says 10% CPU for Devstral; other parameters unstated
  • Harness version: unknown

Variant identity

  • mistralai/Devstral-Small-2-24B-Instruct-2512 · quant Q4_K_M · artifact bartowski/mistralai_Devstral-Small-2-24B-Instruct-2512-GGUF · historical revision unknown · SHA-256 unknown
  • zai-org/GLM-4.7-Flash · quant 4-bit; exact variant unknown · artifact unknown · historical revision unknown · SHA-256 unknown
  • Qwen/Qwen3.5-27B · quant unknown · artifact unknown · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Limitations and corrections

  • Private code; no public rubric/output
  • Different contexts and unspecified quants
  • One task domain; not a general web/agent ranking
  • Manual assistance rather than autonomous agent work

Read The_Paradoxy's source · checked 2026-10-02

M4 64GB comparison: Qwen creates playable Tetris, Devstral and GLM struggleg_rich · 2026-03-19 · Contextual observation

M4 Mac Studio (author wording; exact M4 variant unstated) · 64 GiB reported · Runtime unknown · version unknown · harness OpenCode

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.5-35B-A3B · success: Author reports working game and prefers speed/quality balance.
  • Qwen/Qwen3.5-27B · success: Working game, slower than MoE.
  • zai-org/GLM-4.7-Flash · partial: Game ran but severe bugs made it nearly unplayable.
  • mistralai/Devstral-Small-2-24B-Instruct-2512 · failure: Unplayable game in author's test.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Create Tetris clone in one HTML file
  • configured tokens: unknown
  • output limit tokens: unknown
  • kv cache: unknown
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.5-35B-A3B · quant 8-bit; format unspecified · artifact unknown · historical revision unknown · SHA-256 unknown
  • Qwen/Qwen3.5-27B · quant 8-bit; format unspecified · artifact unknown · historical revision unknown · SHA-256 unknown
  • zai-org/GLM-4.7-Flash · quant 8-bit; format unspecified · artifact unknown · historical revision unknown · SHA-256 unknown
  • mistralai/Devstral-Small-2-24B-Instruct-2512 · quant 8-bit; format unspecified · artifact unknown · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Limitations and corrections

  • One toy task; no outputs/repeated seeds
  • Contrasts with scientific-code Devstral report under different task/hardware/settings
  • No general ranking
  • Qwen3-Coder-Next Q4 mentioned in source but excluded from this candidate-oriented list

Read g_rich's source · checked 2026-10-02

Supporting source: Parent discussion has readable comment

High-precision Qwen3.6 loses compression instructions in a long agent runmilpster · 2026-07-19 · Setup candidate

AMD gfx906 setup; exact GPU/RAM unstated · Memory unknown · llama.cpp · version unknown · harness OpenCode + oh my openagent

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.6-27B · failure: Changed requirement to fixed 40% compression, delegated to wrong agent, lost meaning-preservation rule, sometimes bypassed delegation.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Compress prompts without losing meaning, delegating to a designated agent
  • configured tokens: 105000
  • output limit tokens: unknown
  • kv cache: q8_0 K/V; F16 also tried
  • initial prompt tokens range: 45000–55000
  • alternatives: Up to 256K with F16 also tried
  • temperature: 0.6
  • top p: 0.95
  • top k: 20
  • min p: 0
  • presence penalty: 0.4
  • repetition penalty: 1
  • reasoning preserve: true
  • thinking: true
  • parallel: 1
  • ctx checkpoints: 60
  • batch size: 16384
  • ubatch size: 384
  • speculation: ngram-mod and draft-mtp flags; interaction unresolved
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.6-27B · quant UD-Q8_K_XL; also Q8_0 and down to Q6 · artifact Qwen3.6-27B-UD-Q8_K_XL.gguf · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Limitations and corrections

  • Harness/compaction/reasoning plumbing confounded
  • Suggested fixes not author-confirmed
  • Not code-generation benchmark; relevant to long-agent reliability
  • No hardware capacity or pinned revisions

Read milpster's source · checked 2026-10-02

Qwen builds useful but incomplete Playwright suite on RTX 4090John (johnhringiv) · 2026-06-02 · Setup candidate

Windows 11 workstation · 24 GiB reported · LM Studio · version 0.4.12 build 1; llama.cpp CUDA 12 runtime v2.14.0 · harness OpenCode under WSL2

Cloud plan / local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.6-27B · partial: 74 passing tests, 52 failing; missed coverage, forbidden fixed waits, latent date bugs. Seven operator interventions and four compactions.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Implement Playwright E2E tests for gotflashes Laravel/Livewire application from shared synthesized plan
  • configured tokens: 163840
  • kv cache: q8_0 K/V
  • output limit tokens: unknown
  • temperature: 0.6
  • top p: 0.95
  • top k: 20
  • min p: 0
  • repeat penalty: 1
  • presence penalty: 0
  • thinking: true
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.6-27B · quant UD-Q3_K_XL · artifact unsloth/Qwen3.6-27B-GGUF · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Author-reported measurements

  • agent compute time: 2h39m. Author's run.

Limitations and corrections

  • One repository/run
  • Cloud comparator different harness/context
  • Plan grading involved Claude with author input
  • Passing tests not necessarily correct tests; prompts support inspection, not reproduction
  • Execution is cloud-planned-local-executed: shared synthesis plan came from a Claude Opus plan plus local cases. 74 passing / 52 failing / 14 skipped are author tests, not independent reproduction.

Read John (johnhringiv)'s source · checked 2026-10-02

Supporting source: Published planning prompt

Supporting source: Published implementation prompt

M3 Max 36GB: compact harness achieves seven of eight seeded bug fixesBig_Cycle_6146 · Date unknown · Setup candidate

MacBook Pro M3 Max 14-core CPU / 30-core GPU · 36 GiB reported · oMLX · version 0.6.3rc3 · harness OMP, a Pi fork

Planning provenance unclear. needs task-planning provenance clarification; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.8-27B · partial: Shrinking harness prompt 22.6K to 5.9K yielded author-reported seven of eight bugs fixed. At 21–23K context some requests rejected/empty HTTP-200 streams; Docker memory induced ~1 t/s/aborted streams.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • 80 scored coding-agent runs involving eight seeded bugs in two real repositories and multi-turn sessions
  • configured tokens: 30000
  • server max tokens: 32768
  • output limit tokens: unknown
  • kv cache: TurboQuant KV 3.5-bit
  • memory guard gb: 27
  • concurrency: 2
  • mtp: true
  • temperature: 1
  • top p: 0.95
  • hot cache: 2GB
  • ssd cache: 60GB
  • chunked prefill: false
  • prefill priority: context
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.8-27B · quant AWQ 5.0bpw · artifact True2456/Qwen3.8-27B-AWQ-5.0bpw · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Author-reported measurements

  • logged requests: 1068. Author reported; not independently reproduced.
  • local model hours: about 22. Author reported; not independently reproduced.
  • generation median tokens/s 5k to 10k: 27.9. Author reported; not independently reproduced.
  • generation median tokens/s 25k to 30k: 19.8. Author reported; not independently reproduced.

Limitations and corrections

  • One machine/model/harness
  • Claude produced harness/benchmark/runs/post under human direction
  • Artifacts linked, not independently reproduced
  • Aggressive KV quant/memory constraints differ from Jason Q6 setup
  • Short-prompt peak 32 t/s is not long-session throughput; retain context-binned medians
  • Seven of eight seeded bugs is this test's count, not a community success rate
  • Date remains null. Claude helped create harness/runs/post; local target inference does not establish cloud-free experimental preparation. Seven of eight bugs fixed is within this experiment, not a community success rate.

Read Big_Cycle_6146's source · checked 2026-10-02

Supporting source: Author's claimed presets/scripts/captured requests

MacBook Pro 48GB: MTPLX slowdown resolved by engine patchchettykulkarni · Date unknown · Contextual observation

MacBook Pro; chip not explicitly stated by OP · 48 GiB reported · LM Studio initially, then MTPLX · version unknown · harness OpenCode

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.8-27B · unclear: Latest update reports steady 25+ t/s after engine patch and 25–30 t/s borderline usable; website completion/correctness not established.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Simple ecommerce website with basic edge cases
  • Later long agentic task, details unstated
  • configured tokens: unknown
  • output limit tokens: unknown
  • kv cache: unknown
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.8-27B · quant GGUF 4-bit / MLX 8-bit / MLX 4-bit; final MTPLX 4-bit · artifact unknown · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Author-reported measurements

  • baseline GGUF4 or MLX8 tokens/s: 9–15. Author reported; not independently reproduced.
  • baseline MLX4 tokens/s: 19. Author reported; not independently reproduced.
  • latest MTPLX sustained tokens/s: 25–30. Author reported; not independently reproduced.

Limitations and corrections

  • Throughput report, not demonstrated task success
  • Exact quant publishers/runtime patch/context unknown
  • Updated post date unavailable; latest read 2026-10-02
  • Historical fall ~30 to ~15 then ~3 t/s later reported fixed; not current outcome
  • Do not attach another commenter's M5 Pro chip to OP's machine
  • Date set to null: secondary index date is not a verified primary timestamp. Latest engine patch recovered steady 25+ t/s; task completion remains unclear. Do not infer OP chip from another commenter.

Read chettykulkarni's source · checked 2026-10-02

Supporting source: Linked engine issue; not independent user report

M2 Air 24GB: flight simulator needs roughly 63 hours and remains buggyHyperFoci · Date unknown · Setup candidate

MacBook Air M2 · 24 GiB reported · LM Studio Bionic · version unknown · harness unknown

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.8-27B · partial: 47.8-hour initial run produced title screen with broken start key. Corrective prompt took 15 more hours; result flyable but no plane model and still buggy.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Build relaxing flight simulator in one HTML page
  • configured tokens: 57000
  • output limit tokens: unknown
  • kv cache: unknown
  • precision note: 57K as reported
  • temperature: unknown
  • mtp: unknown
  • reasoning effort: unknown
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.8-27B · quant Q3_K_S · artifact unknown · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Author-reported measurements

  • initial duration hours: 47.8. Author reported; not independently reproduced.
  • repair duration hours: 15. Author reported; not independently reproduced.

Limitations and corrections

  • Extreme low-bit/memory-constrained config, not estimate for M5 48GB Q6
  • No full log or memory-pressure diagnosis
  • Cloud comparisons use different conditions and excluded
  • Date set to null: secondary index date is not a verified primary timestamp. Initial 47.8h plus 15h correction is partial, not successful completion or expected M5 performance.

Read HyperFoci's source · checked 2026-10-02

M2 Max 96GB: faster runtimes, but none of 22 game outputs judged goodex-arman68 · 2026-08-23 · Contextual observation

Mac Studio M2 Max · 96 GiB reported · MTPLX; oMLX + Lightning MTP; llama.cpp +/- MTP/DFlash2; mlx-dspark; vllm-mlx · version llama.cpp baseline b10470; others unspecified · harness Four-phase prompt workflow; agent harness unspecified

Local execution. author-described; not independently reproduced; no cost or cloud-free preparation inferred.

Author-reported outcomes

  • Qwen/Qwen3.8-27B · partial: Most runs delivered code, author judged none of 22 games good; mlx-dspark DFlash2 xhigh used 226K thinking tokens/delivered nothing.

Evaluator: Original author; not independently reproduced. One author investigation; supporting links and variants are not extra independent reports. Affiliations unknown; no independent reproduction.

Reported tasks and settings

  • Plan/implement Bubble Bobble clone in four phases, across runtimes/reasoning settings
  • configured tokens: 262144
  • output limit tokens: 100000
  • kv cache: Unquantized
  • reasoning effort: medium–xhigh
  • sampler: Official Qwen coding sampler
  • chat template: Official Qwen Jinja template
  • llama mtp draft tokens max: 3
  • llama parallel: 1
  • Harness version: unknown

Variant identity

  • Qwen/Qwen3.8-27B · quant 8-bit variants · artifact Unsloth Q8_0 GGUF; mlx-community MLX8; Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality; scottlowry/Qwen3.8-27B-oQ8e-mtp · historical revision unknown · SHA-256 unknown

No verified exact download or hardware recommendation is attached to this observation.

Author-reported measurements

  • total gpu hours: more than 110. Author reported; not independently reproduced.
  • MTPLX medium wall: 1h35m. Author reported; not independently reproduced.
  • MTPLX medium decode tokens/s: 21–24. Author reported; not independently reproduced.
  • llama cpp MTP medium wall: 1h04m. Author reported; not independently reproduced.
  • llama cpp MTP medium decode tokens/s: 18–20. Author reported; not independently reproduced.
  • oMLX MTP medium wall: 1h18m. Author reported; not independently reproduced.

Limitations and corrections

  • Runtime experiment, not independent model-quality comparison
  • Tool-call failures explicitly not tested
  • One machine/task with different artifacts
  • No transfer of general recommendation to M5 48GB
  • Use updated oMLX + MTP rows; baseline mistaken for MTP in discussion
  • Composite throughput/reliability/quality score is not coding accuracy
  • One author investigation spanning several artifacts/runtimes. Updated oMLX rows supersede old MTP labels; do not collapse into a single setup or infer tool-call failures from a task without tool use.

Read ex-arman68's source · checked 2026-10-02

Coverage, freshness & unavailable sources

Manual editorial review, last reviewed 2026-10-02. Source checked dates describe retrieval, not when a test ran. Metadata refresh candidates stay in a private review queue and cannot auto-publish findings. Reports older than a release/runtime change may not apply; source checks older than 30 days are marked stale.

  • Referenced MLX oQ4e-mtp artifact: Anonymous metadata request returned 401. No bypass; exact public download identity remains unverified.
  • Artificial Analysis datasets: Visual inspiration only. No dataset imported or license assumed.

No evidence-backed improvement over your current setup can be inferred without a comparable task/baseline.