Skip to content

Co-ReAct: Rubrics as Step-Level Collaborators

Translation status: The English translation is pending. The section structure is synchronized with the Chinese source.

Motivation: Moving Rubrics from Judge to Guide

Method: Preference Collection, Listwise GRPO, and Inject-Verify-Retry

Collecting Branch-Point Preference Data

Training the Rubric Generator with Listwise GRPO

The Inject-Verify-Retry Loop

Results

Evaluation Setup

Main Results

Ablations

Generalization to Commercial Models

Search-Behavior Analysis

Portability

Case Study

Limitations

Position in the Deep-Research Landscape