Physical Sciences › Computer Science › Information Systems
Cloud Computing and Resource Management
115 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos48 % · 37 artículos
- China34 % · 26 artículos
- Reino Unido10 % · 8 artículos
- Singapur7,8 % · 6 artículos
- India6,5 % · 5 artículos
- Australia5,2 % · 4 artículos
- Alemania5,2 % · 4 artículos
- Italia5,2 % · 4 artículos
Sobre 77 artículos de este tema con al menos un laboratorio localizado. 25 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing
Pakorn Nathong, Kunat Pipatanakul · 2 de octubre de 2026
Instance-billed serverless platforms charge for CPU and memory over the lifetime of a warm instance, making idle inference state a direct serving cost. We present billing-aware neural text-to-speech (TTS) serving on serverless CPUs, optimizing CPU-seconds and GB-seconds rather than throughput or lat…
- ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum
Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi, Pantelis Frangoudis, Schahram Dustdar · 30 de septiembre de 2026
Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pres…
- When Does Learning Beat Heuristics? A Case Study in Kubernetes Scheduler Score Plugins
Wang Xuying, Zhibek Sarypbekova · 22 de septiembre de 2026
Kubernetes scheduler plugins that score candidate nodes are, in production, hand-tuned heuristics. We ask whether a learned scoring function - trained on real placement decisions from a production cluster trace - can match or exceed these heuristics, and if not, why. We implement an external, HTTP-b…
- Real-World Deployment and Performance Characterisation of Fog-Based Deep Learning for Cold-Chain Temperature Prediction over LoRaWAN
Jeremiah Taguta, Jean Frederic Isingizwe Nturambirwe, Clement Nthambazale Nyirenda · 16 de septiembre de 2026
Fresh fruits and vegetables (FFVs) are highly perishable, and cold-chain breaks contribute significantly to global food waste. While Machine Learning (ML) can enable proactive intervention, cloud-based inference faces challenges such as latency and data loss. Fog computing addresses these issues but…
- Leakage-Safe and Scheduler-Aware Machine Learning for Grid Job Runtime Prediction
Ashfaq Ali Shafin, Khandaker Mamun Ahmed · 15 de septiembre de 2026
Accurate job runtime prediction can improve scheduling-aware resource management in grid and distributed computing environments, but prediction models must be evaluated under realistic deployment constraints. This paper revisits CPU burst time prediction on the GWA-T-4 AuverGrid workload trace and r…
- Hyperion: An AI-powered HPC cluster for sciences and humanities research that utilizes ML for predicting job turnaround time
Jun Zhou, Nathan Elgar, Tawnee Benedetto, John Richards, Ming Hu, Greg Wilsbacher, Lawrence Miao, Paul Sagona · 14 de septiembre de 2026
Hyperion is an innovative high-performance computing (HPC) cluster developed for researchers in both science and humanities disciplines at the University of South Carolina (USC). Our approach involved constructing a HPC cluster designed to meet the current research needs while accommodating future e…
- MCRL2: Multi-resource Cross-attention-based Representation Learning-augmented Reinforcement Learning for Cloud Microservice Scheduling
Tiangang Li, Shi Ying, Xiangbo Tian, Chuan Shi, Ding Xiao · 14 de septiembre de 2026
Efficient microservice scheduling is crucial for maintaining load balance across nodes in data centers and ensuring high quality of service. However, achieving this in practice remains challenging due to dynamic resource imbalance under fluctuating workloads, nonlinear coupling across multiple resou…
- Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management
Antonino Vaccarella, Lanpei Li, Vincenzo Lomonaco, Massimo Coppola · 10 de septiembre de 2026
Managing resources across IoT, edge, and cloud layers calls for continuous, context-aware decisions under constraints that rarely stay fixed. Deep reinforcement learning (DRL) handles this class of problems well, and large language models (LLMs) are increasingly used to augment DRL pipelines, yet th…
- Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters
Milos Gravara, Andrija Stanisic, Stefan Nastic · 7 de septiembre de 2026
Compound AI workflows are increasingly used to serve complex AI tasks by coordinating multiple AI models and software components. This approach enables deployment flexibility, as each workflow stage can expose different model variants and resource requirements, but it also expands the deployment cho…
- A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds
Ashir Javeed, Anton Borg, H{\aa}kan Grahn, Lars Lundberg, Dhyey Patel, Sogand Shirinbab · 4 de septiembre de 2026
Accurate cloud resource forecasting is essential for proactive resource provisioning, maintaining Quality of Service (QoS), and reducing operational costs in dynamic cloud environments. The existing forecasting approaches predominantly estimate future CPU workload directly from historical resource t…
- Artificial Intelligence for Energy Optimization in Data Centers
Mohammed Basharath Ullah, Summaiya Unnisa Begum, Mohammed Nadeem Ullah · 4 de septiembre de 2026
Data centers are increasingly optimized by artificial intelligence and, at the same time, increasingly loaded by it. The literature treats these as two unrelated problems: control studies model workload as an exogenous arrival process, while sustainability studies model infrastructure as a fixed mul…
- MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs
Youssef Ennouri, Soonhoi Ha · 3 de septiembre de 2026
Deploying heterogeneous AI models concurrently on a shared GPU introduces resource contention that complicates runtime scheduling. While surrogate models avoid costly online benchmarking, their profiling requirements typically grow combinatorially with the number of co-running models, limiting scala…
- Scalable Multi-GPU Simulation of 3D Multicellular Growth with RNN-Based Workload Balancing
Matvey Moisseyev, Huijing Du, Dandan Zheng, Chi Zhang, Hongfeng Yu · 27 de agosto de 2026
Detailed multicellular growth simulations based on subcellular element models (SEMs) can capture complex tissue development, but their element-level interactions impose substantial computational cost. This work presents a scalable multi-GPU framework for 3D multicellular growth simulation that combi…
- MolGA: Molecular Graph Adaptation with Pre-trained 2D Graph Encoder
Xingtong Yu, Chang Zhou, Xinming Zhang, Yuan Fang · 26 de agosto de 2026
Molecular graph representation learning is widely used in chemical and biomedical research. While pre-trained 2D graph encoders have demonstrated strong performance, they overlook the rich molecular domain knowledge associated with submolecular instances (atoms and bonds). While molecular pre-traini…
- TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
Jie Li, Chenxin Jia, Jinliang Shen, Cunzhuang Liu, Ruiyi Ding, Jianwen Xian, Kang He, Chengru Song · 17 de agosto de 2026
In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below $n^* \approx 15…
- User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling
Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magn\'usson, Praveen Kumar Donta · 13 de agosto de 2026
Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resou…
- Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet
Eliseo Curcio · 13 de agosto de 2026
Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, that slow hardware indiscriminately. We instrumen…
- MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning
Hanye Zhao, Muning Wen, Yong Yu, Weinan Zhang · 11 de agosto de 2026
Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but its required training effort is unknown. Existing approaches combine online loss prediction with adaptive resource allocation, yet commonly treat computation as c…
- ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
Yifan Wu, Yuhan Li, Zhenhua Wang, Ke Chen, Lidan Shou, Zonghao Chen, Liang Lin, Huan Li, Gang Chen · 11 de agosto de 2026
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge. Our analysis of production workloads in Alibaba AnalyticDB ex…
- Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference
Jiming Su, Hantao Hua, Lujia Yin, Yiping Yao, Feng Zhu · 7 de agosto de 2026
In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. Existing multithreaded strategies struggle to match thread resources …
- Multi-tenant Kubernetes Use Cases for AI, Secure Computing and Data Services, and More
Jake Watson, Sadaf R Alam, Christopher Woods, Abdelwahab Kawafi, Thomas Green, Ian Johnson, Ellis Pires, Jessica R. Jones, Utz-Uwe Haus · 4 de agosto de 2026
Kubernetes, as a container orchestration engine, has been widely used in cloud-native ecosystems for several years. In supercomputing ecosystems, especially where bare-metal performance for compute and network devices are considered, the adoption is somewhat limited. However, with the increasing div…
- ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform
Xiaoxiao Jiang, Suyi Li, Sheng Yao, Tianyu Feng, Lingyun Yang, Dapeng Nie, Haoran Yang, Wei Wang · 30 de julio de 2026
Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. Existing platforms typically deploy each workflow as an opaque GPU function, provisioning, placing, and scaling all constituent models in t…
- Right-sizing Recommendations (RSR): Cloud Workload Conformal Prediction for Virtual Machines in Data Center Operations
Mehryar Majd, Feng Cheng, Ali Pahlevan · 29 de julio de 2026
Managing cloud infrastructure efficiently, especially in environments of large cloud providers or hyperscalers, requires optimizing the use of physical resources to minimize costs and maximize performance. Selecting the right virtual machine (VM) sizes is crucial to achieving cost efficiency in thes…
- An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism
Mobina Kashaniyan, Mehrdad Ashtiani, Amirhossein Ghassemi · 20 de julio de 2026
Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging because of dynamic workloads, cold-start latency, and dependencies among functions. We present a dependency-aware autoscaling framework that integrates graph-based bot…
- Quota Marketplace: Dynamic Pricing for Efficient Allocation of ML Training Resources
Balasubramanian Sivan, Renato Paes Leme, Mihai Tiuca, Ian McFarlane, Vasilis Gkatzelis, Nehal Mehta, Soheil Hassas Yeganeh, Vahab Mirrokni, Amin Vahdat · 14 de julio de 2026
The escalating demand for Machine Learning (ML) training resources in recent years has resulted in a substantial gap between the high demand and the available supply. Efficient allocation of these scarce and expensive resources is crucial for organizations to maximize their return on investment. Exi…
Otros asuntos del tema Sistemas de información
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Software Engineering Research853 artículos / 12 meses+218 %
- Information Retrieval and Search Behavior727 artículos / 12 meses+506 %
- Recommender Systems and Techniques600 artículos / 12 meses+154 %
- Expert finding and Q&A systems205 artículos / 12 meses+1175 %
- Information and Cyber Security169 artículos / 12 meses+1500 %
- Blockchain Technology Applications and Security146 artículos / 12 meses+650 %
