Physical Sciences › Computer Science › Computer Networks and Communications
Distributed and Parallel Computing Systems
45 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Últimos artículos
- Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories
Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu, Zhiqiang Zhang, Xiaolin Huang, Jun Zhou · 30 de septiembre de 2026
Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how muc…
- AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization
Pengyu He, Yan Zhang, Ruien Li, Guangwen Yang · 30 de septiembre de 2026
The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronizatio…
- OptiCom : A Unified Framework for State-Conditioned Composition in LLM-Driven Optimization
Chenxing Wei, Sichen Liu, Lizhao Liu, Ningyuan Sun, Chen Bingzhou, Ying He, Bo Jiang, Fei Yu, Yao Shu · 30 de septiembre de 2026
Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Target…
- GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning
Zhengao Li, Shuoqiu Li, Xiaofang Zhang, Yukai Jin, Gokcen Kestor, Yanfu Zhang, Yiming Zeng, Bin Ren, Chuxu Zhang, Shangqian Gao · 29 de septiembre de 2026
Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective…
- PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
Jiaan Zhu, Wei Gao, Youhui Bai, Zewen Jin, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Cheng Li · 29 de septiembre de 2026
Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use als…
- DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving
Junyi Shen, Noppanat Wadlom, Zhengyuan Su, Yao Lu · 28 de septiembre de 2026
Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does not hide it: the key tha…
- When Parallel Drafter Meets Parallel Speculative Decoding
Fuliang Liu, Xue Li, Kun Qian, Zhibin Wang, Wanchun Dou, Wenyuan Yu, Chen Tian · 24 de septiembre de 2026
DSpark-style parallel drafters have made speculative decoding highly effective, yet their draft phase remains serialized on the critical path of every round. Parallel speculative decoding (PSD) overlaps drafting with verification, yet existing methods must guess the accepted prefix and bonus token i…
- WeightBridge: An Efficient Weight Transfer Library for Reinforcement Learning
Xuanlin Jiang, Samuel Hsia, Michael Kuchnik, Zachary DeVito, Minlan Yu, Carole-Jean Wu · 23 de septiembre de 2026
Weight transfer - the propagation of updated parameters from trainers to rollout generators - is becoming an important performance bottleneck in reinforcement learning (RL) systems for LLMs. The central challenge is supporting the diverse trainer and rollout layouts and synchronization requirements …
- Optimizing Denoising Trajectories in dLLMs: A Lightweight Evolutionary Heuristic Approach
Zijian Zhao, Dian Jin, Xialiang Tong, Sen Li, Mingxuan Yuan · 23 de septiembre de 2026
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to conventional Auto-Regressive (AR) Large Language Models (LLMs). By leveraging bidirectional attention and parallel decoding, dLLMs enable more efficient generation. However, they require a carefully designed …
- What Must Survive? Exact Task-Information--State Frontiers for Resource-Sufficient Learning
Ronald Katende · 21 de septiembre de 2026
A system may be compressed before its downstream task is fully known. We ask how much retained state is then necessary and how much can be saved by limited advance task information. For a finite family of linear tasks, a task message is revealed before state formation and the exact task only after…
- Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
Zhenyu Zhang, Jiudong Yang, Zhaowen Tao, Meng Chen · 21 de septiembre de 2026
Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a ligh…
- The Life of a Token: from Words to Bits on the Wire
Davide Avesani (CEDRIC - ROC), Pengwenlong Gu (CEDRIC - ROC), Sotiris Skaperas (Cnam), Stefano Secci (CEDRIC - ROC) · 18 de septiembre de 2026
Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. Behind their ease of use lies a complex process: words become tokens, tokens become vectors, and vectors ultimately give rise to streams of bits that …
- Optimal Model Activation Policies for Inference Networks of Large Language Models
Foivos Charalampakos, Md Ibrahim Ibne Alam, Iordanis Koutsopoulos, Koushik Kar · 16 de septiembre de 2026
Recent advances in large language models (LLMs) have rendered them necessary for Natural Language Processing (NLP) tasks, and their high inference cost motivates the study of cost-performance trade-offs. In practice, several expert LLMs are used in synergy for inference, either in an ensemble mode o…
- LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents
Siddharth Sharma, Nilesh Prasad Pandey, Onat Gungor, Tajana Rosing · 15 de septiembre de 2026
As LLM agents become integrated into increasingly complex workflows, they must continually acquire new capabilities while retaining competence on previously learned tasks. Lifelong agents address this through experience replay, injecting past interactions into the prompt to leverage prior experience…
- HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning
Hongliang Wei (Harbin Institute of Technology, Alibaba Cloud), Xiaobing Tu (Alibaba Cloud), Yinggui Wang (Alibaba Cloud), Zhengxi Liu (Alibaba Cloud), Rongkun Xue (Alibaba Cloud), Jinkui Ren (Alibaba Cloud), Xiantao Zhang (Alibaba Cloud), Debin Zhao (Harbin Institute of Technology), Xiaopeng Fan (Harbin Institute of Technology) · 15 de septiembre de 2026
Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory formats. The same model can perform unevenly across these interfaces, making robustness to harness variation an important objective. A natural approach…
- Towards Optimizing SQL Generation via LLM Routing
Mohammadhossein Malekpour, Nour Shaheen, Foutse Khomh, Amine Mhedhbi · 15 de septiembre de 2026
Text-to-SQL enables users to interact with databases through natural language, simplifying access to structured data. Although highly capable large language models (LLMs) achieve strong accuracy for complex queries, they incur unnecessary latency and dollar cost for simpler ones. In this paper, we i…
- MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving
Tiancheng Zhang, Yulin Chen, Yunfeng Zhao, Shaoyuan Huang, Cheng Zhang, Xiaofei Wang · 15 de septiembre de 2026
The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode instances often suffer from persistent load imbalance, as output le…
- Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference
Tian Jin · 15 de septiembre de 2026
Large language models (LLMs) demonstrate impressive capabilities, but their deployment presents significant efficiency challenges. Autoregressive decoding imposes substantial inference latency and under-utilizes hardware accelerators in low batch size regimes. Discrete diffusion models can generate …
- SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
Zihan Wang, Yuqi Wang, Lei Gong, Cheng Tang, Wenqi Lou, Teng Wang, Chao Wang, Xuehai Zhou · 14 de septiembre de 2026
Mixture-of-Experts (MoE) creates a structural advantage for offloading: only a small fraction of activated experts need to reside in device memory, and if they can be loaded in time for computation, offloading can in principle approach full-load performance, where all model weights reside in device …
- Personalized Execution Time Optimization for Billion-Scale Scheduled Jobs
Yang Liu, Juan Wang, Idris Malik, Zhengxing Chen, Ian Fox, Imani Mufti, Jason Sukumaran, Baokun He, Xiling Sun, Feng Liang · 10 de septiembre de 2026
Scheduled batch jobs are widely used on asynchronous computing platforms to execute enterprise applications such as promotional notifications and candidate pre-computation for recommender systems. Delivering or updating information at the right time is important for user experience and execution imp…
- Parallelism Strategy Chaining for Fast Training Convergence
Minchul Kang, Changyong Shin, Younghun Go, Hyunho Lee, Jinwoo Jeong, Chuck Yoo, Gyeongsik Yang · 9 de septiembre de 2026
Selecting a parallelism strategy - the configuration of data, tensor, and pipeline parallelism degrees together with micro- and global-batch sizes - largely determines the training efficiency of large language models. State-of-the-art methods search for a parallelism strategy offline and select the …
- Speculative Macro Commit for Faster Tool-Using Agents
Zeyu Liu, Souvik Kundu, Peter A. Beerel · 4 de septiembre de 2026
Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and observation can delay subsequent decisions. We introduce \textbf{Speculative Macro Commit} (SMC), a runtime mechanism for a two-tier…
- DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference
Xiaoyang Lu, Belthangady Akash Vi Narayana Pai, Xian-He Sun · 2 de septiembre de 2026
Mixture-of-Experts (MoE) models enable efficient scaling of large language model (LLM) inference but suffer from substantial data-movement overhead when deployed on neural processing unit (NPU)-based systems. Near-Data Processing (NDP) provides a promising way to mitigate this bottleneck via coopera…
- From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov, Danil Taranets, Daniil Dryabin, Mikhail Gashkov, Viktor Zelenkovskiy, Aleksandr Fida, Gleb Alektorov, Nikita Gulyakov, Arthur Babkin, Aleksandr Medvedev, Pavel Gein, Anatolii Potapov · 2 de septiembre de 2026
Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quali…
- Instella-MoE Technical Report
Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra, Yonatan Dukler, Gowtham Ramesh, Jialian Wu, Ximeng Sun, Wen Xie, Chaojun Hou, Vikram Appia, Zhenyu Gu, Zicheng Liu, Emad Barsoum · 2 de septiembre de 2026
In this work, we introduce Instella-MoE, a fully open Mixture-of-Experts (MoE) language model with 16 billion total parameters and 2.8 billion active parameters per token, trained entirely from scratch on AMD Instinct MI300X and MI325X GPUs. Instella-MoE combines a sparsely activated MoE design with…
Otros asuntos del tema Redes informáticas y comunicaciones
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Software System Performance and Reliability395 artículos / 12 meses+400 %
- Constraint Satisfaction and Optimization254 artículos / 12 meses+220 %
- Software-Defined Networks and 5G205 artículos / 12 meses+400 %
- Network Security and Intrusion Detection186 artículos / 12 meses+260 %
- IoT and Edge/Fog Computing150 artículos / 12 meses+175 %
- Caching and Content Delivery139 artículos / 12 meses+1500 %
