Direct Preference Optimization Beyond Chatbots: Fixing Vision Models With Their Own Failures

Everyone uses Direct Preference Optimization for chat alignment, but it turns out to be the ultimate weapon against stubborn OCR degeneration loops....

Feed
October 3, 2026
Direct Preference Optimization Beyond Chatbots: Fixing Vision Models With Their Own Failures


We talk endlessly about aligning chatbots to be helpful and harmless, as if preference tuning only exists to keep AI from sounding rude. That is a massive waste of a powerful tool. Direct preference optimization goes way beyond internet chat. When we built DharmaOCR, we ran headfirst into a classic wall: text degeneration. Models get stuck in infinite repetition loops, spitting out garbage instead of transcriptions. Because token-by-token loss treats local probabilities fine while missing the macro-disaster unfolding right in front of it, supervised fine-tuning hits a hard ceiling here.

The standard training playbook fails because it lacks completion-level judgment. SFT evaluates every single token in isolation, meaning a runaway loop of repeating words never gets flagged as a catastrophic failure of the entire output. It just looks like a string of locally probable guesses. But reality is binary. The text is either a correct extraction of the document, or it is an endless, broken stutter. That clear distinction changes everything about how we should train these systems.

By feeding the model's own failure modes back into the pipeline as rejected pairs, we forced a dramatic shift in behavior without touching the underlying architecture. DPO cuts through the noise. It treats the whole sequence as the unit of truth, punishing the model ruthlessly for stuttering while rewarding clean, accurate transcriptions of Brazilian Portuguese text.

Direct Preference Optimization Beyond Chatbots: Fixing Vision Models With Their Own Failures

The results across open-source families speak for themselves. Every single model tested saw its degeneration rate plummet, dropping by nearly sixty percent on average, with some models plunging by almost ninety percent. We did not need subjective human feedback or hand-crafted conversational datasets. We just looked at where the model broke, gathered its garbage outputs. Also, used DPO to teach it what failure actually looks like. Worth noting. In a sense, stop treating preference tuning like a chatbot luxury. It is a core engineering fix for broken outputs.

If your pipelines are breaking on structural tasks, stop throwing raw compute at another round of standard fine-tuning. Look at your failures. Build your preference pairs. Let the model learn from its worst moments.