Close Menu
    Facebook X (Twitter) Instagram
    • Privacy Policy
    • Terms Of Service
    • Legal Disclaimer
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Facebook X (Twitter) Instagram
    Brief ChainBrief Chain
    • Home
    • Crypto News
      • Bitcoin
      • Ethereum
      • Altcoins
      • Blockchain
      • DeFi
    • AI News
    • Stock News
    • Learn
      • AI for Beginners
      • AI Tips
      • Make Money with AI
    • Reviews
    • Tools
      • Best AI Tools
      • Crypto Market Cap List
      • Stock Market Overview
      • Market Heatmap
    • Contact
    Brief ChainBrief Chain
    Home»AI News»Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models
    Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models
    AI News

    Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models

    July 30, 20267 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email
    synthesia


    Tencent has released AngelSpec, an open-source, torch-native training framework for speculative-decoding draft models. The release covers both autoregressive multi-token prediction (MTP) and the block-parallel DFlash family.

    Most speculative-decoding work searches for one drafter that scores well on an averaged benchmark mixture. Real serving traffic does not look like that mixture. AngelSpec treats workload heterogeneity as a first-class design constraint, and specializes structure, training data, and verification depth around it.

    Why one universal drafter underperforms

    Speculative decoding is lossless. A lightweight drafter proposes several future tokens, and the target model verifies them together in one forward pass using rejection sampling. Acceleration then depends on two things: how many draft tokens get accepted, and how long the complete draft–verify round takes.

    Those two quantities move in opposite directions across domains. In high-entropy open-ended conversation, many continuations are semantically valid. The target may pick any one of them, so acceptance decays quickly with proposal depth. Generating and verifying a long block wastes compute. Autoregressive MTP drafting fits this regime because it proposes a shorter candidate sequence.

    bybit

    Code and mathematical reasoning behave differently. Programming syntax, repeated identifiers, formal expressions, and step-by-step derivations constrain future tokens more strongly. These workloads create longer predictable spans, which is exactly what block-parallel drafting amortizes well.

    AngelSpec therefore ships two complementary drafters, not one compromise. The MTP model is trained on rich, diverse conversation-oriented data. The block-diffusion model is strengthened with code- and mathematics-focused samples.

    The MTP path: Training-Time Test and target-model rollout

    The original Hy3 model is trained with a single MTP layer and no recurrent self-conditioned unrolling. At inference the block can be reused recurrently, but the training objective never prepared it for a long self-generated chain. Errors accumulate with depth, so the second and third draft positions accept at substantially lower rates than the first.

    AngelSpec addresses this train–inference mismatch with a shared-parameter, multi-depth scheme. It retains D logical prediction depths but reuses one physical MTP block. During training that block is autoregressively unrolled for D steps, with each prediction fed into the next invocation. Following the Training-Time Test principle of EAGLE-3, depth k+1 receives the arg max prediction from depth k instead of the ground-truth token. Parameters are shared, but supervision stays depth-specific: the teacher target advances one future position at every depth.

    Two further choices carry most of the gain. First, the target backbone and target language-model head are frozen, and MTP inputs from the backbone are detached. The drafter improves its proposal distribution without touching the distribution that verification must preserve. Second, training uses target-model rollout: responses are generated by the frozen target rather than taken from the original reference corpus. That produces the exact token choices, hidden-state trajectories, and local uncertainty patterns MTP has to approximate at serving time.

    The measured effect is concentrated where it should be. At T = 0, mean acceptance moves from 52.8% to 66.4%, and mean accepted length from 2.58 to 2.99. The first-position rate is nearly preserved at 0.799 → 0.814. The deeper positions carry the delta: p3 climbs from 0.290 to 0.706 on GSM8K, and from 0.387 to 0.757 on HumanEval.

    DFly: hybrid target conditioning and a predecessor-conditioned AR head

    DFly is the block-diffusion architecture. It builds on DFlash with two structural changes.

    Hybrid target-conditioning backbone: DFlash concatenates hidden states from multiple target layers and transforms them with a fully connected layer, producing one shared context feature. That single context is then supplied to every draft layer, which limits layer specialization. DFlare instead learns per-layer fusion weights, giving each draft layer its own view of the target hierarchy, but drops DFlash’s learned cross-layer transformation. DFly composes them: the FC branch establishes a common semantic basis, and the per-layer fusion is applied as a residual refinement on top. The extra branch introduces only D × T scalar weights, and its softmax coefficients can be precomputed after training.

    Predecessor-conditioned autoregressive head: A parallel backbone predicts each block position from the accepted context only. It cannot see which continuation was actually selected at earlier draft positions, which produces suffix acceptance decay. DFly places a small sequential head after the parallel backbone, converting position-wise marginal predictions into prefix-conditioned distributions. The expensive backbone stays fully parallel; only the small head runs left to right.

    Early Performance

    On Qwen3-8B, DFly reaches 5.41 average mean accepted length, against 5.32 for DSpark, 4.57 for DFlash, and 3.24 for MTP. It takes the best result on all five math and code benchmarks. DSpark stays slightly ahead on MT-Bench at 3.77 versus 3.67, which is consistent with DFly being positioned for code and math.

    On Hy3-A21B, the margin is wider. DFly reaches 4.79 against 3.69 for DFlash and 3.00 for MTP — relative gains of 29.8% and 59.7%. It improves every one of the six reported benchmarks.

    The cumulative ablation on Hy3-A21B under greedy decoding traces where that comes from: DFlash backbone 3.77, DFly backbone 4.40, plus Markov head 4.56, and hidden correction 4.60, and code/math data 4.75. The data expansion adds 700K prompts — 500K code from OpenCodeInstruct and OpenCodeReasoning, 200K math from Big-Math. Any prompt sharing a contiguous 16-token span with an evaluation example is removed before response generation.

    Inside the framework

    AngelSpec is built on TorchSpec and extends it in several places. The foundation is disaggregated: inference engines run the frozen target model and stream hidden states through a Mooncake-backed RDMA store directly to distributed training workers. Hidden states are captured inside vLLM worker processes through public vLLM APIs — a speculative hidden-state extraction hook and a custom KV connector — without forking the engine.

    The extensions that matter for this work:

    • TTT rollout unrolled in parallel over the whole sequence, with the causal-prefix-plus-diagonal attention structure reproduced implicitly via compiled FlexAttention and logsumexp merging. Memory stays close to a single causal pass.
    • Long-context training with Ulysses sequence parallelism, validated at context lengths up to 128k tokens. Each local shard carries a D-token halo so depth-shifted supervision stays rank-local.
    • Document-aware sequence packing with three isolation mechanisms — an attention document gate, a depth-shift document gate specific to the MTP path, and document-local position encoding. Cross-document isolation is covered by unit tests verifying zero attention leakage.
    • Evaluation server that periodically runs genuine speculative decoding against the latest checkpoint on dedicated GPUs, reporting mean accepted length and per-position acceptance as measured by the serving engine itself.
    • Pluggable interfaces at three levels: targets (runtime vLLM plugin entry point, no source patch), objectives (composed over a shared base, selected by config), and optimizers (Muon as a drop-in alternative to AdamW).

    Backends are tiered: vLLM is first-class, with SGLang and HuggingFace Transformers supported at community tier.

    Key Takeaways

    • Seven checkpoints ship on Hugging Face and ModelScope, including no-think and high-think DFly variants.
    • AngelSpec trains six draft architectures — DFly, DFlash, DFlare, Eagle3, DSpark, MTP — behind one config-driven pipeline.
    • DFly lifts mean accepted length on Hy3-A21B to 4.79, versus 3.69 for DFlash and 3.00 for MTP.
    • On HY3-295B-A21B with TP=8, DFly-8 delivers a 1.98–2.40× speedup over autoregressive decoding across concurrency 4 to 64.
    • D-cut pushes live-traffic throughput to 981 tok/s at c64, +15.7% over DFly, while giving up 2.8% acceptance.

    Check out the Paper, GitHub Repo, Documentation, Hugging Face Collection and ModelScope Collection. All credit for this research goes to the researchers of this project.

    Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.



    Source link

    Customgpt
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    CryptoExpert
    • Website

    Related Posts

    How a medical database developed at MIT evolved into a global standard of data-sharing | MIT News

    July 29, 2026

    Snowflake launches Cortex AI Gateway to control AI agents and prevent runaway enterprise costs

    July 28, 2026

    How AI is shortening drug discovery timelines in China

    July 27, 2026

    KwaiKAT Team Releases KAT-Coder-V2.5: An Agentic Coding Model Trained on 100,000+ Verifiable Repository Environments

    July 26, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    bybit
    Latest Posts

    Learn AI Filmmaking in 8 Minutes | (Full Step-By-Step Workflow)

    July 30, 2026

    How a $1.9 billion ‘Bitcoin reserve’ ended up in the top corporate rankings without buying a single coin

    July 30, 2026

    South Korea Giants LG CNS and POSCO International Deploy Live Trade Data on Injective Blockchain

    July 30, 2026

    Ethereum Startup EthSystems Targets Institutional Blockchain Privacy

    July 30, 2026

    Wheat Rallying Early on Thursday Amid More Black Sea Strikes

    July 30, 2026
    notion
    LEGAL INFORMATION
    • Privacy Policy
    • Terms Of Service
    • Legal Disclaimer
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Top Insights

    Bitcoin Joins Risk-Asset Relief As PCE Inflation Follows Expectations

    July 30, 2026

    Tokenized Gold Survives DeFi Test as Lending Adoption Lags

    July 30, 2026
    coinbase
    Facebook X (Twitter) Instagram Pinterest
    © 2026 BriefChain.com - All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.