No reproducible way to compare local LLM coding setups
Every discussion of local models for daily coding is a pile of unreproducible anecdotes — quantization, VRAM, context size and agent harness are almost never stated together, so nobody can replicate a reported result.
No must-have criteria recorded.
No bonus criteria.
“I researched this” — self-reported by the poster, not verified by WantBorn. View the linked evidence
Founder-seeded from a real, cited source: original thread
Evidence-based checks from other people — shown as raw counts with the evidence one click away. Not a verification badge; WantBorn doesn't verify identities.
Checks move a quest's status, so they need a real account behind them. Voting stays open without signing in.
Fit scores are the submitter's own claim against the stated must-haves — disputable, never a WantBorn verdict, and nothing but criteria-fit ever touches this ranking.
No solutions suggested yet — know one that fits?
A gap-flagged quest can be claimed by a builder — the claim promotes it into a brief, and “validated” is only ever computed from independent checks plus real test/pledge commitments, never from votes. We're building the Forum in the open, one honest piece at a time.