Jeeves Model Outperforms Competitors with Enhanced Reasoning and Diffusion Drafter
September 29, 2026
A complete reproduction and data pipeline is outlined, detailing scripts for data preparation, training (SFT, CISPO, drafter), and exporting to a fused standalone model that includes a diffusion drafter for serving.
Performance results show Jeeves outpacing Kev-9B and Jev on held-out tests and JevBench public items, with accuracy figures such as 0.889 vs 0.822 and 0.857 for Jev, and 0.935 for Jeeves on JevBench hard/easy public items.
Jeeves is a 9B Jev-like model built from Qwen3.5-9B using LoRA and a pointer head, augmented with a diffusion drafter and CISPO-based training to sharpen reasoning before decision making.
The model supports yes/no, multiple-choice, and score-based questions within a Jev-compatible API, targeting improved performance on out-of-domain tasks and areas where Jev underperforms.
Noted limitations include weaker performance on certain knowledge tasks (MMLU), slower thinking at tail distributions, and calibration nuances; comparisons are drawn against specific items and public tiers.
The repository documents acknowledgments, results, training details, diffusion drafter, data management, and references to related works and prior models.
A diffusion drafter visualizes model reasoning, allowing chain-of-thought to inform final decisions while maintaining efficiency, with block-4 as the default for cost-effective batched decoding.
The training workflow combines SFT (two epochs on eight GPUs with 19,126 questions from 12 public datasets plus synthetic policy data), CISPO (624-step schedule with 9,992 RL questions), and calibration to finalize the checkpoint.
Quickstart and Python SDK guides are provided, including example requests and server setup instructions for local deployment and testing.
Summary based on 1 source
