Cosine similarity between a task title and a group description is not a classifier.

Bench, cases, raw verdicts: backend/bench/categorization in the TODOforAI repo. Total cost ~$0.50.

Setup

TODOforAI boards have groups (SEO, Paid, Frontend…); new tasks get a suggested group. V1: embed title and group name + description, nearest cosine wins. Embedder Qwen/Qwen3-Embedding-4B via DeepInfra, 512 dims, unit-normalized, query instruction “Given a task, retrieve the workstream whose purpose best matches the task”. Abstain if best < 0.50 or margin to runner-up < 0.06.

100% precision on 18 hand cases. Most of my real board in Unsorted.

Judges

Same 138 titles, same 8 groups, same description text, no example tasks — exactly what production sees. Titles written like real todos: short, half Hungarian, 15 nonsense (call mom, asdf), 5 ambiguous.

  • embedding — production, unchanged
  • jevTypeSafe Jev, non-autoregressive decision model on Vercel AI Gateway. One choice question, 8 groups + none. $0.04/MTok in, output free.
  • sonnet — Claude Sonnet 4.6, temp 0, one slug or none
  • opus — Claude Opus 5, same prompt. The truth. My labels are one more opinion.

Trap: Opus 5 reasons before answering; max_tokens: 20 looked like abstaining on 88/138. Give it 400.

Numbers

Judgeagrees with Opuswrong groupmissed138 tasks
embedding77/138160/1128 s
jev130/13853/11224 s
sonnet 4.6131/13870/11257 s

Opus routed 112, called none on 26 (15 nonsense + 11 vague). Embedding: wrong once, missed 54%. Jev: 97% routed, 5 wrong. Sonnet’s 7 wrong are coin flips with Opus. On my real board Jev and Opus agree 21/26.

Why the embedding abstains

Not the margin:

                         agree   wrong   missed
margin 0.06  floor 0.50    76      2       60     ← production
margin 0     floor 0.50    79     41       18
margin 0     floor 0       76     62        0

Drop the gate and Unsorted becomes wrong group. Top-1 is wrong on ~45% of tasks.

bun test flaky on CI          frontend 0.540   development 0.537
call mom                      email 0.540      plg 0.525
asdf                          development 0.663  frontend 0.659

Scores sit in 0.42–0.72; asdf outscores a real task, so no floor finds nonsense. Frontend/Development, Paid/SEO, Enterprise/Paid overlap, so real tasks sit inside any margin. Jev and the LLMs know outcomes: a CI flake is Development because of what fixing it achieves. Similarity can’t say that.

What shipped

Jev routes; embedding is fallback when the gateway is unset or down. Batches of 25, none explicit, accept probability ≥ 0.5. ~0.3 s per task.

Not Sonnet: 57 s vs 24 s and ~100× the price, for a suggestion 400 ms after you stop typing.

Routerper 1,000 tasksper 1M tasks
embedding (~15 tok, $0.01/MTok)$0.0002$0.15
jev (~260 tok catalog per question, $0.04/MTok in, out free)$0.01$10
sonnet 4.6 (~350 in / 10 out)$1.20$1,200
opus 5 (~350 in / ~200 out with reasoning)$7$7,000

List prices, 8 groups. Jev costs ~70× the embedding and routes twice as much.

Jev’s 5 wrong are channel-vs-outcome (sponsor a newsletter, measure signups → email, not paid). Confidence is lower when wrong (0.70 vs 0.94), but floor 0.6 trades 2 wrong for 7 Unsorted. We keep the 5.

Takeaways

  • Precision alone lies. Print recall against something you trust.
  • Score against a strong model, not your own labels.
  • Check the reference’s output budget. Truncated looks cautious.

138 cases is small. It’s enough to see a 54% recall gap; that’s the only claim.