reconcile_gst2b_env

Multi-turn enterprise compliance workflow RL on OpenEnv. Instantiated on Indian GST Input-Tax-Credit reconciliation: ~14M businesses, monthly, still manual. Targeting Scaler AI Labs, Multi-App RL Environment for Enterprise Workflows sub-theme.

16 typed tool verbs · 5 labels · 4-component arithmetic reward · 6 red-team attacks CI-enforced <0.45 · 42 tests · make reproduce bit-identical

👇 Default tab is Tab 4 (Baseline Comparison): composite-reward bar chart with the on-site Day 1 trained Qwen3-4B SFT + P3 GRPO at n=5 mean 0.305. For the 3D fraud-ring visual, select Tab 3 (Circular-Ring Viewer), pick hero seed 9502, click Render. Full README + training evidence + Day 1 charts in the Files tab.

Baseline Comparison: where every policy sits on the reward axis

One chart. All composite-reward totals. Oracle, 6 CI-enforced red-team attacks, the prompted Qwen2.5-3B baseline, and the on-site-trained Qwen3-4B SFT + P3 GRPO (Day 1, A100 SXM4-80GB, n=5 mean 0.305). The dashed line at 0.45 is the red-team attack ceiling enforced in CI. The trained Qwen3-4B bar (solid purple) shows the measured Day 1 result, lifted from 0.280 SFT-only baseline by the P3 length-shaping mitigation. See the LESSONS_LEARNED link below the chart for the FM4 + FM5 research findings behind these numbers.