Physical Sciences › Engineering › Electrical and Electronic Engineering
Power Systems and Technologies
31 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Últimos artículos
- SEED: Self-Speculative Decoding via Implicit Encoder-Decoder
Hankun Lin, Patrick Pynadath, Ruqi Zhang · 1 de octubre de 2026
Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the…
- Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
Gaurav Agarwal, Ashish Garg, Isha Singhal · 1 de octubre de 2026
Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency obj…
- DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification
Rongjian Chen, Minxian Xu, Zhengxin Fang, Kejiang Ye, Chengzhong Xu · 30 de septiembre de 2026
Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable token…
- LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
Hao-Yuan He, Peng-Fei Liu, Si Shen, Ming Li · 30 de septiembre de 2026
Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue …
- Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving
Yi Xu, Ehsan K. Ardestani, Wenyin Fu, Martin Schatz, Krishna Malladi, Zhan Shu, Adnan Aziz, Shobhit Kanaujia, Ajit Mathews, Chunqiang Tang · 24 de septiembre de 2026
As serving capacity demand surpasses that of training, serving efficiency becomes increasingly important. Prefill-decode (P/D) disaggregation improves serving efficiency through specialization and isolation of the two phases. These benefits rest on a static partitioning. Phase demand, however, is no…
- PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models
Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo, Haoxuan Che, Song Guo · 23 de septiembre de 2026
Diffusion language models (dLLMs), such as LLaDA and Dream, have become competitive with autoregressive (AR) LLMs in generation quality while supporting native parallel decoding. A standard acceleration strategy is block-wise decoding, where each forward pass predicts a block of length B and commits…
- TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models
Haibo Hu, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue · 23 de septiembre de 2026
Speculative decoding accelerates large language model inference through collaboration between a lightweight draft model and a target verifier. Existing methods mainly improve the draft side, while the target model is typically kept dense and unchanged. We show that, under domain-specific inference, …
- Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference
Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia, Yutaka Watanobe, Fang Liu · 23 de septiembre de 2026
Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing soft-compression methods either lack query-guided memory selection a…
- H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache
Weifan Jiang, Krishna Teja Chitty-Venkata, Megan Flynn, Reed Meyerson, Zhenting Qi, Tianyu Wu, Eldar Kurtic, Minlan Yu, Alexandre Marques · 22 de septiembre de 2026
Speculative decoding losslessly accelerates large language model inference by having a lightweight draft model predict future tokens for verification by the target model. Recent block diffusion drafters further reduce drafting latency by predicting multiple tokens in parallel. However, existing bloc…
- Acceptance-Aware Draft Model Training for Speculative Decoding
Tianhua Xia, Mugilan Ganesan, Yifei Feng, Haiyu Wang, Maximilian Egger, Sai Qian Zhang · 22 de septiembre de 2026
Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to generate multiple candidate tokens that are verified by the target model in a single forward pass. Its speedup is largely determined by the acceptance length, yet existing draft-model training…
- GDN Tree-Scan: Served Tree Verification for Recurrent-Hybrid Language Models
Zhiyuan Ma · 22 de septiembre de 2026
Tree speculative decoding verifies multiple candidate continuations in one target forward pass. For attention-only transformers, the verifier mainly needs an ancestry mask. Recurrent-hybrid language models break this assumption: a candidate row must also carry the recurrent state that native sequent…
- WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
Hyeongju Ha, Jae-Joon Kim · 22 de septiembre de 2026
Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD),…
- Adapting Tree-Structured Speculative Decoding to DeepSeek-V4 for Efficient Inference
Changxu Liu, Zhaogeng Li · 22 de septiembre de 2026
Repeated execution of the target model during autoregressive decoding is a major source of LLM inference latency. Unlike linear speculation, which follows a single candidate chain, tree-structured speculation retains multiple branches from shared prefixes; under the same budget, this broader coverag…
- Watermarkable Multi-Draft Speculative Sampling via Poisson Processes
Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz G\"und\"uz · 21 de septiembre de 2026
Large language models (LLMs) have achieved state-of-the-art performance across a wide range of tasks, motivating two important aspects of deployment: inference efficiency and output provenance, which can be tackled by speculative sampling and watermarking, respectively. However, recent works have sh…
- SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference
Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T · 21 de septiembre de 2026
Running large language models (LLMs) locally continues to be limited by restrictions of compute and memory on consumer hardware. The popular acceleration technologies, such as quantization, speculative decoding, and adaptive inferencing, offer substantial speed boosts but usually necessitate retrain…
- RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding
Qiao Hu, Yepeng Weng, Bo Zhang, Takehisa Yairi · 21 de septiembre de 2026
Speculative decoding accelerates LLM inference by drafting multiple tokens in parallel, with tree-based methods further improving efficiency through hierarchical structures. Dynamic-tree methods such as EAGLE-3 perform well under greedy decoding via deterministic top-K expansion and global pruning. …
- To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals
Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik, Lior Wolf, Itamar Zimerman · 18 de septiembre de 2026
Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text se…
- LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers
SangLyul Cho, Langqing Cui, Sehoon Kim, Dongsu Han, Insu Han · 16 de septiembre de 2026
Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurrent depths. However, they incur higher decoding latency than standard Transformer models of comparable parameter size because shared weights are access…
- Carryover Drafting: Recycling Rejected States for Speculative Decoding
Jahyun Koo, Sunghyeon Woo, Jaeeun Kil, Jeongtae Lee, Sungjae Lee, Kyomin Jung, Minsub Kim · 15 de septiembre de 2026
Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept several tokens. By construction, verification computes representations for both accepted and rejected tokens. Yet, conventional drafters retain only the re…
- PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling
Weisi Yang, Stephen Xia · 10 de septiembre de 2026
Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency. However, LLMs have heavy computational requirements, which are difficult for resource-constrained mobile a…
- Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding
Fengxiang Bie, Yuqing Jian, Yifan Yu, Zhongzhu Zhou, Zelei Shao, Ben Athiwaratkun, Shuaiwen Leon Song, Chenfeng Xu, Xiaoxia Wu, Tianyi Zhang · 10 de septiembre de 2026
Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development,…
- X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding
Jaeduk Lee, Wan Choi · 10 de septiembre de 2026
This paper investigates collaborative speculative decoding (CoSD), a distributed large language model (LLM) inference framework in which an on-device small language model (SLM) drafts candidate tokens and a server LLM verifies them. Existing CoSD methods assume a shared vocabulary between the SLM an…
- Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
Oszk\'ar Urb\'an, Young D. Kwon, Stylianos I. Venieris, Cecilia Mascolo · 4 de septiembre de 2026
Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static draft-tree shape. Prior work rel…
- Free Pause Tokens
John Langford, Nathan Godey, Giovanni Monea, Yoav Artzi, Harry Dong, Ying Fan, Gustavo de Rosa, Zheng Zhan · 4 de septiembre de 2026
A free pause token gives a language model extra compute to form each next-token prediction (as a pause, or thinking, token does) but carries that compute in a parallel prediction stream over a weight-shared backbone rather than as an extra token in the sequence. It improves next-token prediction by …
- OUTLETS: Output-Length Prediction from Speculative Decoding Backbones
Weihuang Wen, Yingying Liu, Yichuan Liu, Wenqi Zeng, Li Zhou, Chumin Sun, Jie Sun, Tianshu Yu · 2 de septiembre de 2026
The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major challenges for resource provisioning and cluster scheduling. Although output-length prediction can mitigate these issues, existing approaches have key drawbacks: external proxy models add substantial la…
Otros asuntos del tema Ingeniería eléctrica y electrónica
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Advanced Memory and Neural Computing492 artículos / 12 meses+283 %
- Ferroelectric and Negative Capacitance Devices215 artículos / 12 meses+100 %
- Indoor and Outdoor Localization Technologies129 artículos / 12 meses+100 %
- Energy Load and Power Forecasting111 artículos / 12 meses−10 %
- Green IT and Sustainability101 artículos / 12 meses+60 %
- VLSI and FPGA Design Techniques84 artículos / 12 meses+500 %
