Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How is calibration of Jev or Jev-inspired models being evaluated?


For a model that returns probabilities over caller-defined options, I'd check calibration separately from accuracy. On held-out labeled questions, NLL and Brier score measure whether it assigns probability to the right answer; ECE or reliability plots can show whether decisions made at, say, 0.8 confidence are right about 80% of the time. Fit any temperature on a separate calibration split and leave the final test untouched. The split should be by underlying document or scenario, not by individual paraphrase, or near-duplicates leak across it.

I tried this with Bobcat, a Jev-style model I trained on Qwen. A single temperature lowered development ECE from 0.021 to 0.010; its sealed final ECE was 0.012. That is calibration against our own labels, not a measured claim about Jev's calibration. On TypeSafe's 20 published workflow examples, we can compare probability assigned to a reference answer, but that reference is model consensus, not ground truth, and the sample is too small for a general calibration claim. Details and the evaluation code: https://huggingface.co/sanghwa-na/bobcat-1.1




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: