Commit 9eefac7
Align checkpoints/eval to 50/100/150 -> 25/50/75/100, keep all 4
max_steps 150->100, ckpt.interval 50->25 with keep_last=4 (all four
checkpoints retained instead of pruning to the last 3), orchestrator.eval
interval matched to 25 so val40 always runs at each retained checkpoint.
Full 500-task test-set eval at each checkpoint runs standalone after the
weights save, not wired into the live RL loop -- keeps training fast and
keeps the 'test set touched once for the final report' framing honest.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>1 parent b0dafe8 commit 9eefac7
2 files changed
Lines changed: 8 additions & 8 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1 | 1 | | |
2 | | - | |
| 2 | + | |
3 | 3 | | |
4 | 4 | | |
5 | 5 | | |
| |||
41 | 41 | | |
42 | 42 | | |
43 | 43 | | |
44 | | - | |
| 44 | + | |
45 | 45 | | |
46 | 46 | | |
47 | 47 | | |
| |||
65 | 65 | | |
66 | 66 | | |
67 | 67 | | |
68 | | - | |
69 | | - | |
| 68 | + | |
| 69 | + | |
70 | 70 | | |
71 | 71 | | |
72 | 72 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1 | 1 | | |
2 | | - | |
| 2 | + | |
3 | 3 | | |
4 | 4 | | |
5 | 5 | | |
| |||
46 | 46 | | |
47 | 47 | | |
48 | 48 | | |
49 | | - | |
| 49 | + | |
50 | 50 | | |
51 | 51 | | |
52 | 52 | | |
| |||
70 | 70 | | |
71 | 71 | | |
72 | 72 | | |
73 | | - | |
74 | | - | |
| 73 | + | |
| 74 | + | |
75 | 75 | | |
76 | 76 | | |
77 | 77 | | |
| |||
0 commit comments