Physical Sciences › Computer Science › Hardware and Architecture
Parallel Computing and Optimization Techniques
821 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- États-Unis53 % · 250 articles
- Chine36 % · 171 articles
- Royaume-Uni9,1 % · 43 articles
- Allemagne6,5 % · 31 articles
- Inde6,5 % · 31 articles
- Canada5,3 % · 25 articles
- Corée du Sud5,3 % · 25 articles
- R.A.S. chinoise de Hong Kong4,2 % · 20 articles
Sur 474 articles de ce sujet dont au moins un laboratoire est situé. 49 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- QK-Wanda: Coupling Queries and Keys for Unstructured Pruning
Ivan Ilin, Peter Richt\'arik · 2 octobre 2026
Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE recon…
- Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models
Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza · 2 octobre 2026
Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distin…
- MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen · 2 octobre 2026
Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channe…
- TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference
Mukund Agarwalla, Chih-Jen Lin · 2 octobre 2026
Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to ea…
- QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching
Zunhai Su, Yuxuan Sun, Jianchao Tan, Tao Zhang, Ruihan Hu, Yuchen Xie, Xunliang Cai, Ngai Wong · 1 octobre 2026
Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, cha…
- Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space
Yuchen Li, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong · 1 octobre 2026
Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was computed. We introduce Backpropagated Output Momentum (BOM), which instead stores a compact moving average of predictio…
- Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference
Ruiyi Ding, Jie Li, Kang He, Ziyan Liu, Chengru Song, Yuedong Xu, Yuan Cheng · 1 octobre 2026
Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 1…
- Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
Oliver Sieberling, Bharat Runwal, David Jin, Ryan Chin, Rameswar Panda, Yoon Kim · 1 octobre 2026
Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends th…
- JustFit: Just-in-Time State Management for Local LLM Serving
Yuhua Chen · 1 octobre 2026
Local agents need memory for model execution and working history. We present JustFit, an MLX runtime that coordinates their overlapping allocations: KVExec executes and checkpoints four-bit KV with bounded workspace, PhaseSwap loads phase-dependent components, and StateTrans preserves history across…
- Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi · 1 octobre 2026
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In t…
- Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
Sietse Schelpe · 1 octobre 2026
A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work t…
- Switching Linear Attention
Hyun Dong Lee, Xavier Gonzalez, Nicolas Zucchet, E. Kelly Buchanan, Emily B. Fox, Scott W. Linderman · 1 octobre 2026
Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequ…
- OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading
Jingyuan Xiao (Tianjin University, Tianjin, China), Jiayue Wang (Tianjin University, Tianjin, China), Yitao Hu (Tianjin University, Tianjin, China), Xinning Wang (Tianjin University, Tianjin, China), Shi Chen (Tianjin University, Tianjin, China), Ziqi Gong (Tianjin University, Tianjin, China), Zhengchao Wang (Tianjin University, Tianjin, China), Guotao Yang (Tianjin University, Tianjin, China), Sheng Chen (Tianjin University, Tianjin, China), Keqiu Li (Tianjin University, Tianjin, China) · 30 septembre 2026
Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a na…
- When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression
Haeyong Kang, Chang D. Yoo · 30 septembre 2026
Training-free KV-cache compression methods such as SnapKV, H2O, and PyramidKV evict tokens at the end of prefill, aiming to preserve the attention mass that future queries are expected to use -optimizing what to keep. We show that this objective fails in two distinct ways. (1) Compensation: restorin…
- Improving Test-Time Scaling with Adaptive Looped Transformers
Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding, Yu Wang · 30 septembre 2026
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs gro…
- Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
Liming Liu, Mingze Wang, Tuo Zhao · 30 septembre 2026
As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typical serving, prompt prefill runs in parallel and is compute-bound, whereas autoregressive decode is sequential and memory-traffic-bound. Conventional width or dep…
- LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica · 30 septembre 2026
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeate…
- RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust
Eugene Hauptmann, Nataliya Kosmyna · 30 septembre 2026
Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combine…
- SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving
Chuan Liu, Shuoming Zhang, Zhicheng Li, Qianqi Sun, Ruiyuan Xu, Qiuchu Yu, Xiyu Shi, Huimin Cui, Jiacheng Zhao · 30 septembre 2026
No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice…
- Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis · 30 septembre 2026
Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes i…
- IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence
Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, Wei Liu, Yinggan Xu, Yunxiang Lu, Zai Zheng, Zhirui Xie, Zhongyang Che, Ziyan Tang, Zuoxiang Zhao, Jian Yao · 30 septembre 2026
We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verificatio…
- AI as a Compiler: Compiling Triton kernels without the Triton compiler
Fran\c{c}ois Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini · 30 septembre 2026
Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA …
- SMat-Attention: Structured Long-Context Sequence Modeling
Emile Anand, Abdullah Ateyeh, Archer Wang, Marin Solja\v{c}i\'c · 30 septembre 2026
Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can c…
- Validating Memory-Optimal Transformer Kernels on Real Hardware: From Formal Derivation to Measured Performance Across Two HPC Clusters
Lenore M. Mullin, Gaetan Hains · 29 septembre 2026
We validate memory-optimal cost functions for transformer kernels derived via the Mathematics of Arrays (MoA). Companion Papers I-IV formally derive kernels for attention forward, backward, fused forward+backward, decode, and the complete block (RMSNorm, gated MLP) as a hardware-independent specific…
- The Extender: A Log-Structured Transformer
Jakob Eriksson (UIC) · 29 septembre 2026
We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each lay…
