Build Your Own Decision Model: A 1.7B LLM Scores 59% on CommonsenseQA

Decision models like Jev answer multiple-choice questions in a single forward pass by constraining output to fixed options. I replicated this with Qwen3-1.7B, achieving 59.4% accuracy on CommonsenseQA, and improved it to 62.4% with a quick finetune. However, the model was severely overconfident: predictions with 90-100% confidence were only 70% accurate. Temperature scaling (T=3.8) reduced expected calibration error from 20.1% to 2.9%. A GitHub repo provides scripts for building, evaluating, finetuning, and calibrating your own model.

The model tends to be extremely overconfident in the 0.9 - 1.0 bin but it's only correct 70% of the time.

More from this day

2026-10-10