Repository navigation
Conversation
Add Qwen hybrid architecture, bounded RL statistics, CPU-offloaded full-backbone training, frozen N-gram lifecycle, distributed checkpoint integrity, and SFT/RL/OPD validation coverage.
ruijieguo
marked this pull request as draft
September 27, 2026 06:18
qsa_indexer_temperature defaulted to 1.0, but the HF reference divides the summed relu(qk) indexer score by sqrt(index_head_dim) before both top-k selection and the KL target. Top-k is scale-invariant, so this only mattered for indexer_distillation_loss: every shipped SFT/RL/OPD config left the default in place, making the trained indexer's KL target ~sqrt(128)=11.3x sharper than the reference calibration -- the design spec explicitly flagged this as a value that must not be silently conflated. qsa_indexer_temperature is now Optional[float] = None; Qwen4ExpConfig resolves an unset value to qsa.default_indexer_temperature(indexer_head_dim) in __post_init__ (falling back to 1.0 when QSA is not configured at all), and still honors and validates an explicit override. Verified on .181: the new config-level regression (test_qwen4_exp_qsa_indexer_temperature.py) passes under the real Megatron/TE environment, and the full local CPU-runnable qwen4_exp suite is unaffected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Resolve 16 conflicts against alibaba/ROLL v0.4.0 (581046a), keeping the Qwen3.8-Flash-Next adaptations on top of upstream's new designs: - offload_states_patch.py: port onto upstream's backend-based offload store (put_tensors/get_tensors/delete_tensors). Re-express CPU-master model offload, HybridDeviceOptimizer phase handling, grad-buffer release/restore and checkpoint-failure frame clearing as layers over the backend API. offload_adam_states tolerates the legacy (optimizer, device) call. - megatron_strategy.py: load_states/offload_states keep include_frozen_parameters and CpuMasterWeightProvider on upstream's backend flow; streaming checkpoint save strategy retained (bounded memory for full-backbone checkpoints). - vllm: keep dev-build-safe compat dispatch (call_maybe_await, signature-driven vllm_utils, ray_distributed_executor), adopt upstream headless/mp paths plus Qwen3.8 GDN backend and frozen-ngram loading hooks. - user_defined_rollout_loop.py: re-inject per-sample sampling_params provenance into upstream's _default_token_postprocess; keep module-level wrapper. - protocol.py / upload_utils.py / context_managers.py / model_update.py: keep both sides' additions (timeout sentinel + strip_multi_modal_for_reward, _link_or_copy + register(), NCCL suspend/resume, weight_provider + convert_kwargs). - sft_worker.py: upstream dispatch (DP_MP_COMPUTE, prefetch) plus torch.no_grad. - tests: chained-optimizer fixture gains the backend attributes upstream requires. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- vllm: dev builds expose the modern Ray executor but no legacy "<pkg>.v1" package; bind CustomRayDistributedExecutor as the v1 executor directly instead of importing a module that does not exist. - sft_worker: honor upstream's auto_resume alongside resume_from_checkpoint (latest checkpoint first, explicit path as fallback), matching BaseWorker. - tests: supply the new required config fields (auto_resume, reward_system_config) and attach an explicit offload backend to the real optimizers the gated Hybrid phase tests build. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… store - offload_states_patch: rebind Hybrid float16 shards onto the parked DDP buffer so old CUDA allocations are released; in CPU-master mode park full-shaped zero-storage views so shard validation still passes while the model is parked; checkpoint_grad_buffer_offload hands each leaf the backend and key prefix it needs. - megatron_strategy: load_checkpoint attaches the offload backend and key prefix to the optimizer before releasing gradients. - tests: resume/LoRA-phase fixtures supply the backend, strategy_args and the now-required parametrized arguments. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- asset lifecycle: the fake config exposes use_value_head and the fake converter streams tensors through iter_mca_state_dict_from_hf, the API the per-parameter loader now calls. - RL statistics: inner_forward_step reads strategy.model.config for router replay; stub it and the R2 record probe. - beam search: stub roll.third_party.vllm.compat.call_maybe_await, which vllm_strategy now imports. Two failures remain in these files and are not caused by the merge: the hybrid-Adam padding check in test_checkpoint_write_memory fails identically on the pre-merge branch, and the beam-search multimodal DataProto ndarray assertion fails identically on pristine upstream v0.4.0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ream merge Upstream v0.4.0 dropped the per-expert rank division in the grouped row and column LoRA layers, but this branch's checkpoints, validators and export tests all rely on expert adapters carrying rank r // topk (6 for r=64, topk=10) while shared modules keep r. Restored at both sites. On 2 H800 ranks the real grouped-LoRA Megatron/TE test failed 4 of 6 cases without it and passes 6 of 6 with it. The 8-GPU LoRA SFT validation had also failed its native rank check on the merged tree: expert tensors were saved at rank 64 where the verified checkpoints use 6. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ability Upstream's gate compares the release tuple, which rejects development builds reporting 0.1.devN even though they expose the modern Ray executor API. Judge dev builds by capability, matching roll/third_party/vllm/__init__.py, and add a regression test; released versions below 0.16 are still rejected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With the TransferQueue backend now on by default, materialized batches carry NonTensorStack columns that the dump process cannot json-encode, so every rollout dump file came out empty (TypeError in the writer process). Convert anything exposing tolist(), not just ndarrays. The regression test fails the same way on pristine upstream v0.4.0 and passes with the change. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Upstream v0.4.0 added TrainingArguments.moe_permute_fusion defaulting to False. update_with_args writes any non-None argument over the model config, so every Flash-Next run silently lost Qwen4ExpConfig's True default. That swaps TE's deterministic FP32 top-k combine for an unfused BF16 scatter_add, and the GRPO frozen-reference sentinel then differed by up to 0.35 nats across data-parallel ranks (bitwise equal before the merge), failing the run's own check. Default the argument to None so the model default survives; an explicit True or False still overrides. The new test fails on the merged default and passes with this change. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
vllm_strategy asserts torch.distributed is not initialized, which fails when earlier tests in the same pytest process leave a process group behind. Patch the check in the new test, as the sweep showed it failing only in that order. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Upstream v0.4.0 added ModelConfig.use_value_head and post_converter now reads mca_config.use_value_head on both the weight and the metadata path. Our two hand-built SimpleNamespace config stubs predate that field, so collecting either file raised AttributeError before any assertion ran. Add the field with the upstream default. On an 8xH800 node this takes the LoRA conversion suite to 12 passed and streaming export to 8 passed plus 2 CPU skips, 10 passed with the native vLLM gate on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Upstream v0.4.0 added two eager re-exports to the rewards package, and one of them, Geo3kRewardWorker, imports mathruler at module scope. mathruler is not installed by the base image and is declared in no requirements file, so importing roll.pipeline.rlvr.rewards raises ModuleNotFoundError. base_pipeline.create_clusters_parallel pre-imports each cluster's worker by path, which imports that package, so every RLVR pipeline now aborts with "Failed to pre-import worker class" no matter which reward worker it configured. Three Flash-Next validation runs died this way while asking only for MultipleChoiceBoxedRuleRewardWorker, which has no such dependency. Resolve the geo3k export lazily through PEP 562 __getattr__. The attribute stays importable and __all__ still advertises it, so a run that wants the worker gets it, and a run that wants mathruler still sees its own ModuleNotFoundError. The new test covers both directions; it errors on collection before this change and passes after. It also makes upstream's own gsm8k reward test collectable, which exposes a separate pre-existing failure in that file that this change does not address. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
Qwen3.8-Flash-Next needs its GR residual streams, PLE external N-gram table, QSA attention and hybrid GDN layers preserved throughout training, checkpointing and rollout. This draft adds the model adapter and the ROLL/runtime paths for frozen-table backbone training with CPU optimizer offload, LoRA SFT/RL/OPD, save/resume, and teacher/reference handling.
The latest fixes preserve native GR FP32 FMA, BF16 materialization and padded down/injection GEMM geometry, including the different normalization reduction used after residual combination. Bare projections share one GEMM; wrapped or locally hooked projections retain their normal adapter and synchronization callbacks with the same base GEMM geometry. Analytic backwards retain trainability and activation recomputation. Qwen's router now keeps the checkpoint's model-dtype projection unless the user explicitly selects FP32/FP64; unrelated model defaults are unchanged. SFT also honors explicit step limits without changing epoch-based defaults when the limit is omitted. Flash-Next LoRA workers select a single K partition for dense Triton shrink to avoid nondeterministic FP32 atomic accumulation. This is scoped to eligible model workers and preserves native cached configurations, other kernels, and wrappers across worker reinitialization; it does not enable global batch invariance.
Validation
Fresh H800 validation on September 27:
e6de032, a fresh eight-GPU 8K LoRA SFT process restored checkpoint-3, completed updates 4–5, and saved checkpoint-5. Scheduler count reached six; eight Adam shards, eight worker RNG files, pipeline RNG and adapter structure passed inspection. Training losses were 1.293615 / 1.289751 and heldout losses 2.205300 / 2.121889. This is current-source lifecycle evidence, not convergence or exact continuation equality.e6de032, full-backbone RL completed two real optimizer updates on all eight ranks, followed by native weight transfer and generation. Independent checks verified changed parameter samples on every rank, 32 original N-gram table hashes, an unchanged frozen reference at all three versions, and changed actor outputs after each update. Both steps had nonzero reward advantages. This probe writes no checkpoint and does not establish convergence or long-context numerical parity.e6de032, full-backbone OPD completed two real updates on all eight ranks with changed parameter samples, 32 original N-gram table checks, unchanged trained-teacher sentinels across versions 0/1/2, changing student sentinels, and final native transfer/generation. Both teacher-derived advantage checks passed, with 3,486 and 3,608 nonzero effective tokens. No checkpoint was written; this is lifecycle evidence.37e5756, deterministic dense LoRA shrink boundary tests and CUDA regressions passed all 16 cases on H800. Related local CPU coverage passed 35 cases (the installed-vLLM case was skipped locally). Independent review covered sequential worker reinitialization and composition with foreign wrappers; both findings were reproduced and fixed. The earlier ordinary-WorkerV1 lifecycle passed the same active shrink path; the final source lifecycle is queued behind the sustained backbone RL run.Earlier revisions have eight-GPU SFT/RLVR/OPD lifecycle receipts and LoRA native export/reload/sleep/wake receipts. They do not establish acceptance of the latest numerical changes.
Rebase onto upstream v0.4.0
This branch now contains alibaba/ROLL v0.4.0 (
581046a) as a merge, so the PR no longer conflicts withmain. 16 files conflicted; the large one wasroll/third_party/megatron/offload_states_patch.py, which is re-expressed on top of upstream's new backend-based offload store (put_tensors/get_tensors/delete_tensors) while keeping CPU-master model offload,HybridDeviceOptimizerphase handling, gradient-buffer release and checkpoint-failure frame clearing.Re-running the real 8-GPU paths on the merged tree found regressions that unit tests and import checks did not. All are fixed in this branch:
r // moe_router_topkfrom the grouped row/column LoRA layers. This branch's checkpoints, validators and export tests rely on expert adapters at rankr // topk(6 for r=64, topk=10); without it the 8-GPU LoRA SFT run saved rank-64 expert tensors and failed its own rank check. The 2-rank grouped-LoRA Megatron/TE test went from 4 failed to 6 passed. Upstream's HF-exportrank_patternstill usesr // topk, so the removal looks inconsistent on their side too.moe_permute_fusiondefault. The newTrainingArguments.moe_permute_fusion(defaultFalse) overwroteQwen4ExpConfig'sTruethroughupdate_with_args, replacing TE's deterministic FP32 top-k combine with an unfused BF16 scatter-add. The GRPO frozen-reference sentinel then differed by up to 0.35 nats across data-parallel ranks (bitwise equal before the merge). The argument now defaults toNone; an explicitTrueorFalsestill overrides. The regression test fails on the old default and passes with the change.0.1.devN) that expose the modern Ray executor API are judged by capability instead of the release tuple.NonTensorStackcolumns made every dump file empty (aTypeErrorin the writer subprocess; the same regression test fails on pristine v0.4.0). Fixed by converting anything withtolist().SFTWorkerhonors the newauto_resume.Config defaults that changed upstream and now apply to Flash-Next runs (the validation configs pin the first one):
use_sequence_packingnow defaults toTrue(wasFalse) and is forced off only for non-Megatron strategies or when dynamic batching conflicts. The Qwen4 decoder, QSA attention and bounded RL token statistics all reject packed sequences withNotImplementedError, so Flash-Next Megatron roles needuse_sequence_packing: false(read from the code; no packed run was attempted). The validation configs already set it.pure_opd_pipeline_typenow defaults toNoneand resolves to"rlvr"in pure-OPD mode; a diff of the fully resolved LoRA OPD config against the pre-merge branch shows only new fields, the packing defaults of roles the run does not instantiate, and the transfer backend below.transfer_backendnow defaults to TransferQueue (the previous default was no backend).Controls: against a pristine v0.4.0 tree, 10 of 20 upstream offload-state variants and the multimodal beam-search test fail identically (environment or upstream issues, not regressions here), and the hybrid-Adam padding test fails identically on the pre-merge branch.
Post-merge validation (8 x H800)
GRPO (20 updates, LoRA student, 8 GPUs) on
67c67e8, seed 70: passed. The driver finished (run.exit=0) and the independent verifier passed: 20 optimizer updates on each of 8 ranks, 31 distinct effective prompt groups, parameter changes observed on every rank (at least 975,566 changed sampled elements per rank), frozen reference and frozen backbone unchanged across all 21 weight versions, dynamic adapter transfer verified, and 344 of 344 native N-gram table checks equal to the original checkpoint. This run writes no checkpoint, so it is lifecycle evidence only, not checkpoint or numerical acceptance.Why seed 70. The two earlier 20-step attempts with the default seed 42 had identical data order (it is
random.Random(seed + epoch)). Both ran all 20 steps but failed the driver's end-of-run evidence check: at step 10 all 32 sampled responses hit the 128-token cap,max_len_maskremoved them, the batch had zero effective tokens, and the pipeline's existing skip guard (unchanged from upstream) skipped that optimizer update, while the check expects an update on every step. Replaying the driver's own validator on those results with only that one sequence check made skip-aware passed everything else, but neither run is counted. The validator and the independent verifier were left unchanged. Only the validation driver's configured seed and its matching seed assertion changed: two gitignoredoutput/files, recorded asseed_overridein the job manifest and not part of this PR. Seed 70 was chosen by simulating data order with per-prompt stop rates from four earlier runs (the simulation reproduced the real step-to-prompt mapping in 20 of 20 steps); it had the lowest rough estimate, about 7% against 33% for seed 42 (from smoothed per-prompt stop rates), of some step drawing only never-finishing prompts.LoRA SFT (6 updates, 8 GPUs, EP8, 8K sequences) on
67c67e8:run.exit=0, all sevencheck_sft.pychecks passed. Per-step train loss 2.40 → 1.78, final heldout loss 2.03. Checkpoints at steps 3 and 5 with 8 optimizer shards, 8 worker RNG files and 110,848 adapter tensors; expert adapters are rank 6 and shared adapters rank 64, matching the previously verified checkpoints. (The same 6-step job also passed on9a72bc6.)Env-gated Megatron CUDA suites, re-run on H800 against the exact
67c67e8source tree: hybrid phase, CPU optimizer, CPU-master weight streaming and grad staging (41 passed); strategy checkpoint resume, LoRA frozen-phase offload, DDP initialization and gradient clipping (12 passed); and the 2-rank grouped expert LoRA test (6 passed on each rank). The current head4725516differs from67c67e8in test files and one runtime file,roll/pipeline/rlvr/rewards/__init__.py(the lazy export below).CPU sweep of
tests/andmcore_adapter/testson the67c67e8tree: the only failures absent from pristine v0.4.0 were the hybrid-Adam padding test (fails identically on the pre-merge branch) and the new expert-parallel dev-build test, which depended on test order and is fixed in73d95da.Upstream v0.4.0 defect, fixed in
4725516.roll/pipeline/rlvr/rewards/__init__.pyeagerly re-exportsGeo3kRewardWorker, which importsmathrulerat module scope, andmathruleris declared in no requirements file.base_pipeline.create_clusters_parallelpre-imports each cluster's worker by path, so in an environment withoutmathruler(our validation container has no network) RLVR jobs aborted with "Failed to pre-import worker class", including jobs that only configureMultipleChoiceBoxedRuleRewardWorker. It reproduces on pristine v0.4.0. The fix is a PEP 562 lazy export that keeps the name in__all__; the newtests/pipeline/test_rewards_lazy_exports.pyerrors at collection before and gives 5 passed / 1 skipped after. Side effect: upstream'stests/pipeline/test_gsm8k_math_reward_worker.pybecomes collectable and shows a separate failure that predates this PR (compute_rewards() missing 1 required positional argument, becausegetattr(..., '__wrapped__', ...)returns the unbound function). That file is untouched here.759463eadds theuse_value_headfield that v0.4.0's converter now reads to two converter test stubs.SFT cold resume on
67c67e8: bit-exact. A 2-step job resumed from checkpoint-3 at step 4 and finished step 5. The comparison reportsfailure_count: 0over 408,806 tensor leaves, 7,475,287,104 elements, 297,936 optimizer storage items, 8 adapter payloads and 8 worker RNG files, and loss, grad_norm and sequence-length statistics equal the continuous run. Its own limitation: frozen scope is checked by flags and counts, not by full frozen-tensor checksums.HF adapter export of the
67c67e8SFT checkpoint-5 (CPU only): 148,808 tensors and 1,747,148,800 elements, identical to the earlier export; shared adapters rank 64, expert adapters rank 6. CUDA was not initialised, and no reload or inference was run on this export.LoRA OPD (20 updates, rank-64 student, 8 GPUs) on
67c67e8: passed.run.exit=0and the completion record reports 20 optimizer updates on each of 8 ranks, 52 distinct effective prompt groups, about 1.00 million changed sampled elements per rank, frozen reference and frozen-backbone sentinel unchanged across 21 weight versions, and 344 of 344 native N-gram table checks. The teacher is the rebuiltopd-teacher-v3, not the teacher of the earlier verified run, and its metadata hashes differ. A first attempt was cut off at 14 of 20 steps by a 6 h timeout and is not counted; the counted run is a full 20-step cold start, because the driver asserts it cannot resume. This is lifecycle evidence:full_opd_acceptance,full_rl_acceptanceandcheckpoint_validationare all false.Eight-rank suites:
expert_parallel1 passed,tp4_ep8_loss9 passed,qwen3_5_hybrid6 passed / 2 failed. The failure that was examined (test_torch_gdn_backend_bypasses_conv_and_delta_kernels) monkeypatches an attribute the convolution-patched Megatron checkout used for the run does not define, so it depends on the checkout.Full-backbone RL and OPD (8 GPUs) on
67c67e8: not passed, cause not determined. Both completed step 0 with valid metrics (RLtrain_infer_diff_token_mean0.00144,train_infer_kl0.00117; OPD -0.00015 and 0.00146;ratio_mean1.0 and 32 samples in both). At the start of step 1 Ray's memory monitor killed a worker: the node was at 2003.91 of 2015.50 GiB for RL and 2003.98 of 2015.50 GiB for OPD, against a 0.97 threshold (1955.0 GiB). Each of the 8ActorWorkerprocesses held about 183.45 GiB, which matches the per-worker RSS of the pre-merge run. That run, same RL config file, completed 20 updates one6de032, so this is neither shown to be a capacity limit of the node nor shown to be a merge regression. The OPD run used a different teacher from the pre-merge one, so it has no clean comparison. One untested candidate: idle Ray workers left in the container by earlier sessions, roughly 48 to 60 GiB by an RSS estimate that overcounts shared pages. They were cleaned up afterwards and the two jobs have not been re-run.Remaining acceptance gates
Full-backbone RL and OPD have not passed on the merged tree (previous section); checkpoint, save/resume numerical parity and quality gates are separate and not covered by the lifecycle runs above.
Full-model native probability parity is still failing at atol=rtol=0.02. Fixing native prefill chunks to 2,048 tokens produced exact same-prefix probabilities across 2K/8K, exact aligned 2,049-token controls and exact repeated requests in all three domains. This supplies a stable scheduling-controlled native reference, but actor/native parity still fails. Global batch-invariant mode rejects GDN_ATTN. The tolerance has not been relaxed, and full numerical acceptance remains outstanding.
The deterministic LoRA shrink changes reduction scheduling. A kernel-only microbenchmark on an idle H800 over eight shapes puts single-partition shrink at 1.01 to 2.10 times the native selection (geometric mean 1.22, 42 to 89 microseconds per call); workload-level throughput impact has not been measured.
The historical full-backbone continuation comparison reported 864 FP32 alignment-gap differences. A later padding scan does not establish whole-state equality. Full-backbone state equality and export acceptance remain outstanding.
External N-gram validation checks index/header geometry, tensor type/dtype/shape and byte coverage with bounded per-storage memory. It does not hash every payload byte; referenced model assets must remain immutable during a run. N-gram tables remain frozen in this scope.