PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
Abstract
Posterior Label Correction DPO improves preference optimization by routing noisy pairwise labels into clean, flipped, or tied cases using calibrated policy-reference margins.
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.
Community
Preference datasets often contain incorrect preference directions or weak/ambiguous pairs. PLC-DPO introduces a latent clean, flip, or tie state for each pair and uses the calibrated policy–reference margin to infer posterior-like routing weights. These determine whether training reinforces, reverses, or suppresses a strong directional update.
PLC-DPO combines forward, reverse-direction, and tie-regularizing losses using these weights, with EMA calibration, warm-up, and confidence gating for stable routing. Rather than simply filtering suspicious pairs, it attempts to correct their training signal while reusing the log-probabilities already required by DPO, without an auxiliary model or additional supervision. Across 57 dataset–model–benchmark cells, PLC-DPO achieves the best mean win rate against DPO (60.5%).
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization (2026)
- Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization (2026)
- Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation (2026)
- AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection (2026)
- Stop Replacing Noise with Noise: Two-Source Reliability Assessment for Label Correction and Sample Reweighting in Label-Noise Learning (2026)
- Co-Evolving LLM Evaluators and Policies via DynamicRubric (2026)
- Disentangling Optimization Scale from Preference Scale in DPO (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.30597 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
