Skip to content

feat: add experimental Qwen3.8-Flash-Next post-training support - #502

Draft
ruijieguo wants to merge 45 commits into
alibaba:mainfrom
ruijieguo:feat/qwen38-flash-next-training
Draft

ruijieguo wants to merge 45 commits into
alibaba:mainfrom
ruijieguo:feat/qwen38-flash-next-training

Conversation

@ruijieguo

@ruijieguo ruijieguo commented Sep 27, 2026 •

Copy link
Copy Markdown

What this changes

Qwen3.8-Flash-Next needs its GR residual streams, PLE external N-gram table, QSA attention and hybrid GDN layers preserved throughout training, checkpointing and rollout. This draft adds the model adapter and the ROLL/runtime paths for frozen-table backbone training with CPU optimizer offload, LoRA SFT/RL/OPD, save/resume, and teacher/reference handling.

The latest fixes preserve native GR FP32 FMA, BF16 materialization and padded down/injection GEMM geometry, including the different normalization reduction used after residual combination. Bare projections share one GEMM; wrapped or locally hooked projections retain their normal adapter and synchronization callbacks with the same base GEMM geometry. Analytic backwards retain trainability and activation recomputation. Qwen's router now keeps the checkpoint's model-dtype projection unless the user explicitly selects FP32/FP64; unrelated model defaults are unchanged. SFT also honors explicit step limits without changing epoch-based defaults when the limit is omitted. Flash-Next LoRA workers select a single K partition for dense Triton shrink to avoid nondeterministic FP32 atomic accumulation. This is scoped to eligible model workers and preserves native cached configurations, other kernels, and wrappers across worker reinitialization; it does not enable global batch invariance.

Validation

Fresh H800 validation on September 27:

  • 58 GR tests passed on H800, including exact native BF16 forward comparisons, composed normalization boundaries, shared weights, noncontiguous inputs, adapter active/disabled/merged states, module hooks and independent FP64 complete-mix derivative checks with both GR projections adapted. Projection regressions first failed nine cases, and hook regressions first failed three cases. Both are fixed.
  • Six router tests passed, including real Megatron gate forward/backward comparisons, explicit overrides and unchanged defaults for other models. Complete tiny-model forward/backward and activation-recomputation tests also passed (two cases).
  • Eight-GPU 8K LoRA SFT completed two updates and checkpoint inspection. A fresh process restored that checkpoint with the GR changes, completed two more updates, and saved again. The scheduler reached four updates; eight optimizer shards, eight worker RNG files, pipeline RNG and adapter payload structure passed inspection. This continuation predates the router-default change and is a lifecycle regression, not a convergence result.
  • At commit e6de032, a fresh eight-GPU 8K LoRA SFT process restored checkpoint-3, completed updates 4–5, and saved checkpoint-5. Scheduler count reached six; eight Adam shards, eight worker RNG files, pipeline RNG and adapter structure passed inspection. Training losses were 1.293615 / 1.289751 and heldout losses 2.205300 / 2.121889. This is current-source lifecycle evidence, not convergence or exact continuation equality.
  • At e6de032, full-backbone RL completed two real optimizer updates on all eight ranks, followed by native weight transfer and generation. Independent checks verified changed parameter samples on every rank, 32 original N-gram table hashes, an unchanged frozen reference at all three versions, and changed actor outputs after each update. Both steps had nonzero reward advantages. This probe writes no checkpoint and does not establish convergence or long-context numerical parity.
  • At e6de032, full-backbone OPD completed two real updates on all eight ranks with changed parameter samples, 32 original N-gram table checks, unchanged trained-teacher sentinels across versions 0/1/2, changing student sentinels, and final native transfer/generation. Both teacher-derived advantage checks passed, with 3,486 and 3,608 nonzero effective tokens. No checkpoint was written; this is lifecycle evidence.
  • The current SFT checkpoint-5 exported 148,808 finite BF16 adapter tensors through the public converter. Native file loading, EP filtering and packing matched all tensor values on eight EP ranks. Native file-adapter GPU validation also passed on eight ranks and three Chinese/English/code cases after the dense LoRA shrink repair: repeated execution, base return, adapter removal/reload, and level-1 sleep/wake all produced exactly identical probabilities and generated tokens. Both an isolated causal control and the ordinary ROLL WorkerV1 passed independent receipt checks. Before the repair, identical adapter requests changed probabilities by up to 2.94 while the base remained exact. Four real CUDA primitive cases reproduced all 128 drifting native trials and became exact with single-partition shrink; the public regression also checks an independent FP64 reference.
  • At the preceding GR+router source, eight-GPU TP2/EP8 full-backbone capacity validation completed two 8K forward/backward/CPU Adam updates with finite gradients and observed parameter changes. Maximum allocated memory was 72.05 GiB per GPU. This uses synthetic tokens, freezes the external N-gram table and writes no checkpoint; it is not convergence or checkpoint acceptance.
  • A 48-layer TP4/EP8 probe completed 12 finite Chinese/English/code cases at 2K/8K across both router precisions, with exact data-parallel replica losses. The two precision variants share one model load and retain natural routing.
  • Matching TP8/EP8 topology, native custom-all-reduce-disabled inference, native Triton GDN inference, and the corrected GR projection each completed full-model controls. None resolves the outstanding full-model probability gate. The native reference uses FlashInfer GDN by default; the Triton control is explicitly separate.
  • At 37e5756, deterministic dense LoRA shrink boundary tests and CUDA regressions passed all 16 cases on H800. Related local CPU coverage passed 35 cases (the installed-vLLM case was skipped locally). Independent review covered sequential worker reinitialization and composition with foreign wrappers; both findings were reproduced and fixed. The earlier ordinary-WorkerV1 lifecycle passed the same active shrink path; the final source lifecycle is queued behind the sustained backbone RL run.
  • Focused CPU fallback tests, whitespace checks and independent code review completed. The projection review identified a hook-bypass defect that is covered by the final regression and fixed.

Earlier revisions have eight-GPU SFT/RLVR/OPD lifecycle receipts and LoRA native export/reload/sleep/wake receipts. They do not establish acceptance of the latest numerical changes.

Rebase onto upstream v0.4.0

This branch now contains alibaba/ROLL v0.4.0 (581046a) as a merge, so the PR no longer conflicts with main. 16 files conflicted; the large one was roll/third_party/megatron/offload_states_patch.py, which is re-expressed on top of upstream's new backend-based offload store (put_tensors / get_tensors / delete_tensors) while keeping CPU-master model offload, HybridDeviceOptimizer phase handling, gradient-buffer release and checkpoint-failure frame clearing.

Re-running the real 8-GPU paths on the merged tree found regressions that unit tests and import checks did not. All are fixed in this branch:

  • Grouped expert LoRA rank. Upstream removed r // moe_router_topk from the grouped row/column LoRA layers. This branch's checkpoints, validators and export tests rely on expert adapters at rank r // topk (6 for r=64, topk=10); without it the 8-GPU LoRA SFT run saved rank-64 expert tensors and failed its own rank check. The 2-rank grouped-LoRA Megatron/TE test went from 4 failed to 6 passed. Upstream's HF-export rank_pattern still uses r // topk, so the removal looks inconsistent on their side too.
  • moe_permute_fusion default. The new TrainingArguments.moe_permute_fusion (default False) overwrote Qwen4ExpConfig's True through update_with_args, replacing TE's deterministic FP32 top-k combine with an unfused BF16 scatter-add. The GRPO frozen-reference sentinel then differed by up to 0.35 nats across data-parallel ranks (bitwise equal before the merge). The argument now defaults to None; an explicit True or False still overrides. The regression test fails on the old default and passes with the change.
  • CPU-master offload. Parked parameters keep full-shaped zero-storage views so shard validation holds while the model is parked, Hybrid float16 shards are rebound off the old CUDA allocation, and checkpoint loading hands the offload backend to the leaf optimizers.
  • vLLM expert-parallel gate. Development builds (0.1.devN) that expose the modern Ray executor API are judged by capability instead of the release tuple.
  • Rollout dumps. With the TransferQueue backend now the default, NonTensorStack columns made every dump file empty (a TypeError in the writer subprocess; the same regression test fails on pristine v0.4.0). Fixed by converting anything with tolist().
  • SFT resume. SFTWorker honors the new auto_resume.
  • Tests. Stale fixtures updated for v0.4.0 interfaces.

Config defaults that changed upstream and now apply to Flash-Next runs (the validation configs pin the first one):

  • use_sequence_packing now defaults to True (was False) and is forced off only for non-Megatron strategies or when dynamic batching conflicts. The Qwen4 decoder, QSA attention and bounded RL token statistics all reject packed sequences with NotImplementedError, so Flash-Next Megatron roles need use_sequence_packing: false (read from the code; no packed run was attempted). The validation configs already set it.
  • pure_opd_pipeline_type now defaults to None and resolves to "rlvr" in pure-OPD mode; a diff of the fully resolved LoRA OPD config against the pre-merge branch shows only new fields, the packing defaults of roles the run does not instantiate, and the transfer backend below.
  • transfer_backend now defaults to TransferQueue (the previous default was no backend).

Controls: against a pristine v0.4.0 tree, 10 of 20 upstream offload-state variants and the multimodal beam-search test fail identically (environment or upstream issues, not regressions here), and the hybrid-Adam padding test fails identically on the pre-merge branch.

Post-merge validation (8 x H800)

  • GRPO (20 updates, LoRA student, 8 GPUs) on 67c67e8, seed 70: passed. The driver finished (run.exit=0) and the independent verifier passed: 20 optimizer updates on each of 8 ranks, 31 distinct effective prompt groups, parameter changes observed on every rank (at least 975,566 changed sampled elements per rank), frozen reference and frozen backbone unchanged across all 21 weight versions, dynamic adapter transfer verified, and 344 of 344 native N-gram table checks equal to the original checkpoint. This run writes no checkpoint, so it is lifecycle evidence only, not checkpoint or numerical acceptance.

  • Why seed 70. The two earlier 20-step attempts with the default seed 42 had identical data order (it is random.Random(seed + epoch)). Both ran all 20 steps but failed the driver's end-of-run evidence check: at step 10 all 32 sampled responses hit the 128-token cap, max_len_mask removed them, the batch had zero effective tokens, and the pipeline's existing skip guard (unchanged from upstream) skipped that optimizer update, while the check expects an update on every step. Replaying the driver's own validator on those results with only that one sequence check made skip-aware passed everything else, but neither run is counted. The validator and the independent verifier were left unchanged. Only the validation driver's configured seed and its matching seed assertion changed: two gitignored output/ files, recorded as seed_override in the job manifest and not part of this PR. Seed 70 was chosen by simulating data order with per-prompt stop rates from four earlier runs (the simulation reproduced the real step-to-prompt mapping in 20 of 20 steps); it had the lowest rough estimate, about 7% against 33% for seed 42 (from smoothed per-prompt stop rates), of some step drawing only never-finishing prompts.

  • LoRA SFT (6 updates, 8 GPUs, EP8, 8K sequences) on 67c67e8: run.exit=0, all seven check_sft.py checks passed. Per-step train loss 2.40 → 1.78, final heldout loss 2.03. Checkpoints at steps 3 and 5 with 8 optimizer shards, 8 worker RNG files and 110,848 adapter tensors; expert adapters are rank 6 and shared adapters rank 64, matching the previously verified checkpoints. (The same 6-step job also passed on 9a72bc6.)

  • Env-gated Megatron CUDA suites, re-run on H800 against the exact 67c67e8 source tree: hybrid phase, CPU optimizer, CPU-master weight streaming and grad staging (41 passed); strategy checkpoint resume, LoRA frozen-phase offload, DDP initialization and gradient clipping (12 passed); and the 2-rank grouped expert LoRA test (6 passed on each rank). The current head 4725516 differs from 67c67e8 in test files and one runtime file, roll/pipeline/rlvr/rewards/__init__.py (the lazy export below).

  • CPU sweep of tests/ and mcore_adapter/tests on the 67c67e8 tree: the only failures absent from pristine v0.4.0 were the hybrid-Adam padding test (fails identically on the pre-merge branch) and the new expert-parallel dev-build test, which depended on test order and is fixed in 73d95da.

  • Upstream v0.4.0 defect, fixed in 4725516. roll/pipeline/rlvr/rewards/__init__.py eagerly re-exports Geo3kRewardWorker, which imports mathruler at module scope, and mathruler is declared in no requirements file. base_pipeline.create_clusters_parallel pre-imports each cluster's worker by path, so in an environment without mathruler (our validation container has no network) RLVR jobs aborted with "Failed to pre-import worker class", including jobs that only configure MultipleChoiceBoxedRuleRewardWorker. It reproduces on pristine v0.4.0. The fix is a PEP 562 lazy export that keeps the name in __all__; the new tests/pipeline/test_rewards_lazy_exports.py errors at collection before and gives 5 passed / 1 skipped after. Side effect: upstream's tests/pipeline/test_gsm8k_math_reward_worker.py becomes collectable and shows a separate failure that predates this PR (compute_rewards() missing 1 required positional argument, because getattr(..., '__wrapped__', ...) returns the unbound function). That file is untouched here. 759463e adds the use_value_head field that v0.4.0's converter now reads to two converter test stubs.

  • SFT cold resume on 67c67e8: bit-exact. A 2-step job resumed from checkpoint-3 at step 4 and finished step 5. The comparison reports failure_count: 0 over 408,806 tensor leaves, 7,475,287,104 elements, 297,936 optimizer storage items, 8 adapter payloads and 8 worker RNG files, and loss, grad_norm and sequence-length statistics equal the continuous run. Its own limitation: frozen scope is checked by flags and counts, not by full frozen-tensor checksums.

  • HF adapter export of the 67c67e8 SFT checkpoint-5 (CPU only): 148,808 tensors and 1,747,148,800 elements, identical to the earlier export; shared adapters rank 64, expert adapters rank 6. CUDA was not initialised, and no reload or inference was run on this export.

  • LoRA OPD (20 updates, rank-64 student, 8 GPUs) on 67c67e8: passed. run.exit=0 and the completion record reports 20 optimizer updates on each of 8 ranks, 52 distinct effective prompt groups, about 1.00 million changed sampled elements per rank, frozen reference and frozen-backbone sentinel unchanged across 21 weight versions, and 344 of 344 native N-gram table checks. The teacher is the rebuilt opd-teacher-v3, not the teacher of the earlier verified run, and its metadata hashes differ. A first attempt was cut off at 14 of 20 steps by a 6 h timeout and is not counted; the counted run is a full 20-step cold start, because the driver asserts it cannot resume. This is lifecycle evidence: full_opd_acceptance, full_rl_acceptance and checkpoint_validation are all false.

  • Eight-rank suites: expert_parallel 1 passed, tp4_ep8_loss 9 passed, qwen3_5_hybrid 6 passed / 2 failed. The failure that was examined (test_torch_gdn_backend_bypasses_conv_and_delta_kernels) monkeypatches an attribute the convolution-patched Megatron checkout used for the run does not define, so it depends on the checkout.

  • Full-backbone RL and OPD (8 GPUs) on 67c67e8: not passed, cause not determined. Both completed step 0 with valid metrics (RL train_infer_diff_token_mean 0.00144, train_infer_kl 0.00117; OPD -0.00015 and 0.00146; ratio_mean 1.0 and 32 samples in both). At the start of step 1 Ray's memory monitor killed a worker: the node was at 2003.91 of 2015.50 GiB for RL and 2003.98 of 2015.50 GiB for OPD, against a 0.97 threshold (1955.0 GiB). Each of the 8 ActorWorker processes held about 183.45 GiB, which matches the per-worker RSS of the pre-merge run. That run, same RL config file, completed 20 updates on e6de032, so this is neither shown to be a capacity limit of the node nor shown to be a merge regression. The OPD run used a different teacher from the pre-merge one, so it has no clean comparison. One untested candidate: idle Ray workers left in the container by earlier sessions, roughly 48 to 60 GiB by an RSS estimate that overcounts shared pages. They were cleaned up afterwards and the two jobs have not been re-run.

Remaining acceptance gates

Full-backbone RL and OPD have not passed on the merged tree (previous section); checkpoint, save/resume numerical parity and quality gates are separate and not covered by the lifecycle runs above.

Full-model native probability parity is still failing at atol=rtol=0.02. Fixing native prefill chunks to 2,048 tokens produced exact same-prefix probabilities across 2K/8K, exact aligned 2,049-token controls and exact repeated requests in all three domains. This supplies a stable scheduling-controlled native reference, but actor/native parity still fails. Global batch-invariant mode rejects GDN_ATTN. The tolerance has not been relaxed, and full numerical acceptance remains outstanding.

The deterministic LoRA shrink changes reduction scheduling. A kernel-only microbenchmark on an idle H800 over eight shapes puts single-partition shrink at 1.01 to 2.10 times the native selection (geometric mean 1.22, 42 to 89 microseconds per call); workload-level throughput impact has not been measured.

The historical full-backbone continuation comparison reported 864 FP32 alignment-gap differences. A later padding scan does not establish whole-state equality. Full-backbone state equality and export acceptance remain outstanding.

External N-gram validation checks index/header geometry, tensor type/dtype/shape and byte coverage with bounded per-storage memory. It does not hash every payload byte; referenced model assets must remain immutable during a run. N-gram tables remain frozen in this scope.

Add Qwen hybrid architecture, bounded RL statistics, CPU-offloaded full-backbone training, frozen N-gram lifecycle, distributed checkpoint integrity, and SFT/RL/OPD validation coverage.
@ruijieguo ruijieguo changed the title feat: add Qwen3.8-Flash-Next post-training support feat: add experimental Qwen3.8-Flash-Next post-training support Sep 27, 2026
@ruijieguo
ruijieguo marked this pull request as draft September 27, 2026 06:18
ruijieguo and others added 14 commits September 27, 2026 15:00
qsa_indexer_temperature defaulted to 1.0, but the HF reference divides the
summed relu(qk) indexer score by sqrt(index_head_dim) before both top-k
selection and the KL target. Top-k is scale-invariant, so this only mattered
for indexer_distillation_loss: every shipped SFT/RL/OPD config left the
default in place, making the trained indexer's KL target ~sqrt(128)=11.3x
sharper than the reference calibration -- the design spec explicitly flagged
this as a value that must not be silently conflated.

qsa_indexer_temperature is now Optional[float] = None; Qwen4ExpConfig
resolves an unset value to qsa.default_indexer_temperature(indexer_head_dim)
in __post_init__ (falling back to 1.0 when QSA is not configured at all), and
still honors and validates an explicit override. Verified on .181: the new
config-level regression (test_qwen4_exp_qsa_indexer_temperature.py) passes
under the real Megatron/TE environment, and the full local CPU-runnable
qwen4_exp suite is unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Resolve 16 conflicts against alibaba/ROLL v0.4.0 (581046a), keeping the
Qwen3.8-Flash-Next adaptations on top of upstream's new designs:

- offload_states_patch.py: port onto upstream's backend-based offload store
  (put_tensors/get_tensors/delete_tensors). Re-express CPU-master model
  offload, HybridDeviceOptimizer phase handling, grad-buffer release/restore
  and checkpoint-failure frame clearing as layers over the backend API.
  offload_adam_states tolerates the legacy (optimizer, device) call.
- megatron_strategy.py: load_states/offload_states keep include_frozen_parameters
  and CpuMasterWeightProvider on upstream's backend flow; streaming checkpoint
  save strategy retained (bounded memory for full-backbone checkpoints).
- vllm: keep dev-build-safe compat dispatch (call_maybe_await, signature-driven
  vllm_utils, ray_distributed_executor), adopt upstream headless/mp paths plus
  Qwen3.8 GDN backend and frozen-ngram loading hooks.
- user_defined_rollout_loop.py: re-inject per-sample sampling_params provenance
  into upstream's _default_token_postprocess; keep module-level wrapper.
- protocol.py / upload_utils.py / context_managers.py / model_update.py: keep
  both sides' additions (timeout sentinel + strip_multi_modal_for_reward,
  _link_or_copy + register(), NCCL suspend/resume, weight_provider + convert_kwargs).
- sft_worker.py: upstream dispatch (DP_MP_COMPUTE, prefetch) plus torch.no_grad.
- tests: chained-optimizer fixture gains the backend attributes upstream requires.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- vllm: dev builds expose the modern Ray executor but no legacy "<pkg>.v1"
  package; bind CustomRayDistributedExecutor as the v1 executor directly
  instead of importing a module that does not exist.
- sft_worker: honor upstream's auto_resume alongside resume_from_checkpoint
  (latest checkpoint first, explicit path as fallback), matching BaseWorker.
- tests: supply the new required config fields (auto_resume,
  reward_system_config) and attach an explicit offload backend to the real
  optimizers the gated Hybrid phase tests build.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… store

- offload_states_patch: rebind Hybrid float16 shards onto the parked DDP
  buffer so old CUDA allocations are released; in CPU-master mode park
  full-shaped zero-storage views so shard validation still passes while the
  model is parked; checkpoint_grad_buffer_offload hands each leaf the
  backend and key prefix it needs.
- megatron_strategy: load_checkpoint attaches the offload backend and key
  prefix to the optimizer before releasing gradients.
- tests: resume/LoRA-phase fixtures supply the backend, strategy_args and
  the now-required parametrized arguments.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- asset lifecycle: the fake config exposes use_value_head and the fake
  converter streams tensors through iter_mca_state_dict_from_hf, the API the
  per-parameter loader now calls.
- RL statistics: inner_forward_step reads strategy.model.config for router
  replay; stub it and the R2 record probe.
- beam search: stub roll.third_party.vllm.compat.call_maybe_await, which
  vllm_strategy now imports.

Two failures remain in these files and are not caused by the merge: the
hybrid-Adam padding check in test_checkpoint_write_memory fails identically
on the pre-merge branch, and the beam-search multimodal DataProto ndarray
assertion fails identically on pristine upstream v0.4.0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ream merge

Upstream v0.4.0 dropped the per-expert rank division in the grouped row and
column LoRA layers, but this branch's checkpoints, validators and export
tests all rely on expert adapters carrying rank r // topk (6 for r=64,
topk=10) while shared modules keep r.

Restored at both sites. On 2 H800 ranks the real grouped-LoRA Megatron/TE
test failed 4 of 6 cases without it and passes 6 of 6 with it. The 8-GPU LoRA
SFT validation had also failed its native rank check on the merged tree:
expert tensors were saved at rank 64 where the verified checkpoints use 6.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ability

Upstream's gate compares the release tuple, which rejects development builds
reporting 0.1.devN even though they expose the modern Ray executor API. Judge
dev builds by capability, matching roll/third_party/vllm/__init__.py, and add
a regression test; released versions below 0.16 are still rejected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With the TransferQueue backend now on by default, materialized batches carry
NonTensorStack columns that the dump process cannot json-encode, so every
rollout dump file came out empty (TypeError in the writer process). Convert
anything exposing tolist(), not just ndarrays. The regression test fails the
same way on pristine upstream v0.4.0 and passes with the change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Upstream v0.4.0 added TrainingArguments.moe_permute_fusion defaulting to
False. update_with_args writes any non-None argument over the model config,
so every Flash-Next run silently lost Qwen4ExpConfig's True default. That
swaps TE's deterministic FP32 top-k combine for an unfused BF16 scatter_add,
and the GRPO frozen-reference sentinel then differed by up to 0.35 nats
across data-parallel ranks (bitwise equal before the merge), failing the
run's own check. Default the argument to None so the model default survives;
an explicit True or False still overrides. The new test fails on the merged
default and passes with this change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
vllm_strategy asserts torch.distributed is not initialized, which fails when
earlier tests in the same pytest process leave a process group behind. Patch
the check in the new test, as the sweep showed it failing only in that order.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Upstream v0.4.0 added ModelConfig.use_value_head and post_converter now
reads mca_config.use_value_head on both the weight and the metadata path.
Our two hand-built SimpleNamespace config stubs predate that field, so
collecting either file raised AttributeError before any assertion ran.

Add the field with the upstream default. On an 8xH800 node this takes the
LoRA conversion suite to 12 passed and streaming export to 8 passed plus 2
CPU skips, 10 passed with the native vLLM gate on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Upstream v0.4.0 added two eager re-exports to the rewards package, and one of
them, Geo3kRewardWorker, imports mathruler at module scope. mathruler is not
installed by the base image and is declared in no requirements file, so
importing roll.pipeline.rlvr.rewards raises ModuleNotFoundError.

base_pipeline.create_clusters_parallel pre-imports each cluster's worker by
path, which imports that package, so every RLVR pipeline now aborts with
"Failed to pre-import worker class" no matter which reward worker it
configured. Three Flash-Next validation runs died this way while asking only
for MultipleChoiceBoxedRuleRewardWorker, which has no such dependency.

Resolve the geo3k export lazily through PEP 562 __getattr__. The attribute
stays importable and __all__ still advertises it, so a run that wants the
worker gets it, and a run that wants mathruler still sees its own
ModuleNotFoundError. The new test covers both directions; it errors on
collection before this change and passes after. It also makes upstream's own
gsm8k reward test collectable, which exposes a separate pre-existing failure
in that file that this change does not address.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant