ReportGem ReportGem

Academic paper

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards

Authors: Yunhao Wang, Binghong Wu, Zhenyu Huang, Jiacheng Shi, Shuo Huang, Tinghao Yu, Feng ZhangPublished: 2026-08-01Paper ID: 2608.00536Category: cs.CVLicense: CC BY 4.0

Abstract

Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it remains hard to optimize in the high-accuracy regime because such rewards become weakly discriminative: near-correct outputs receive very similar scores, providing limited learning signal for hard cases. We propose Step-Aware Annealing (SAA), a plug-and-play reward sharpening mechanism that progressively increases reward curvature during training, amplifying subtle quality differences among high-scoring samples while preserving stability in early learning. Built on SAA, we introduce DocPO, a document policy optimization framework with element-specific, reference-based rewards anchored by edit-distance signals: normalized string edit distance (NED) for text, tree edit distance similarity (TEDS) for tables, and a hybrid Rubric+edit reward for formulas. Experiments on OmniDocBench and DocElemHard show that SAA consistently improves GRPO-style RL across document elements over non-annealed rewards, without requiring additional human supervision for reward construction.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader