Physical Sciences › Computer Science › Artificial Intelligence
Neural Networks and Applications
245 papiers indexés
Les réseaux de neurones artificiels explorent des mécanismes d’apprentissage et d’optimisation à travers des approches variées, comme l’amélioration des données par des techniques telles que le Stochastic Weight Averaging ou l’étude de modèles minimaux pour expliquer des phénomènes comme les lois d’échelle. Certains travaux se concentrent sur la structure même de ces réseaux, en analysant leur courbure, leur symétrie ou leurs fonctions d’activation, tandis que d’autres examinent des architectures spécifiques comme les Probabilistic Circuits ou les dual-encoder heads pour en réduire la complexité. D’autres encore abordent des questions théoriques ou pratiques, comme la certification des performances, l’adaptation à des orientations inédites d’objets ou l’intégration de contraintes physiques dans l’apprentissage.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- États-Unis28 % · 18 articles
- Chine14 % · 9 articles
- France9,2 % · 6 articles
- Allemagne9,2 % · 6 articles
- Royaume-Uni7,7 % · 5 articles
- Espagne7,7 % · 5 articles
- Suède6,2 % · 4 articles
- Canada6,2 % · 4 articles
Sur 65 articles de ce sujet dont au moins un laboratoire est situé. 28 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning
Syon Mansur, Joel Zylberberg · 2 octobre 2026
End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes ach…
- Misalignment of Low-Loss Regions Causes Grokking
Yongding Tian, Zaid Al-Ars, Maksim Kitsak, Peter Hofstee · 2 octobre 2026
Grokking refers to the delayed emergence of validation-set generalization after a model has already overfit the training set. Although first observed in small algorithmic tasks trained with transformers, its underlying mechanism remains unsettled. In this work, we develop an analysis framework based…
- Beyond Affine Transformations: A Soft Dominance Layer for Coordinate-Wise Neural Computation
Mariano Rivera · 2 octobre 2026
This paper presents a preliminary study of an alternative to the affine transformation underlying conventional neural-network layers. In the proposed Soft Dominance Layer, each output unit compares input coordinates with a learnable reference vector and aggregates smooth inequality responses. A sigm…
- The Life Cycle of a Massive Activation: Stochastic Birth, Weight-Decay-Driven Growth, and Competitive Consolidation
S. Aaron McClendon, Jorge Gallego-Feliciano, Antonios Saravanos · 2 octobre 2026
Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled intervent…
- Per-Node Activation Function Evolution in Indirectly Encoded Substrates: Solvability, Limits, and Emergent Diversity
Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel · 2 octobre 2026
Biological neurons achieve computational diversity through specialized types: tonic, bursting, adapting, and fast-spiking cells coexist within the same circuit. Artificial neural networks, by contrast, apply a single activation function uniformly to all nodes, which limits what they can represent. W…
- SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials
Amirhosein Azarpour, Seyyed Moein Kazemi · 2 octobre 2026
Kolmogorov-Arnold Networks (KANs) represent a paradigmatic shift in deep learning by replacing fixed node activations with learnable univariate functions on edges, offering enhanced interpretability and parameter efficiency. While recent polynomial-based KAN variants have addressed the computational…
- The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State · 1 octobre 2026
Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states f…
- Patch Rebirth: Fast and Transferable Model Inversion of Vision Transformers
Seongsoo Heo, Dong-Wan Choi · 1 octobre 2026
Model inversion is a widely adopted technique in data-free learning that reconstructs synthetic inputs from a pretrained model through iterative optimization, without access to original training data. Unfortunately, its application to state-of-the-art Vision Transformers (ViTs) poses a major computa…
- Effective Does Not Mean Useful: Conditional Functional Substitutability for Redundancy and Scaling in Transformers
Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Xinyong Cai, Juncheng Bu, Lan Yu, Tinghe Zhang · 1 octobre 2026
Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We view redundancy as an input-conditioned, dynamic relation: intermediat…
- How Local Mixing Encodes Relative Position in Global NoPE Attention
Cutter Dawes, Nick Alonso, Tom Figliolia, Beren Millidge · 30 septembre 2026
The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encod…
- Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
Zahra Khodagholi, Niloofar Yousefi · 30 septembre 2026
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting t…
- Task-Relevant Null-Space Residuals for Non-Injective Neural Mappings
Bizu Feng, Zhimu Yang, Shuming Wang, Yuan Cheng, Shaode Yu, Xiaojun Qian, Zixin Hu · 30 septembre 2026
Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-…
- Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
Soutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron, Nadav Dym · 29 septembre 2026
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical founda…
- From Distributions to Stochastic Processes: Neural Approximation of Measure-Valued Maps
Yichen Wang, Ziyi Wang, Wenlian Lu, Chenghuang Shen, Jianfeng Liu, Zhengdong Xiao, Longjiu Luo, Qianrong Wang · 29 septembre 2026
Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learning and extend it to mappings between stoc…
- \L{}ukasiewicz Neural Networks Extended: Residual Architectures and Crystallization Strategies for Interpretable Rule Extraction
Carlos Leandro · 29 septembre 2026
A feed-forward neural network whose weights are integers and whose activation is the truncated identity implements, neuron by neuron, the connectives of \L{}ukasiewicz many-valued logic. This exact correspondence --- established theoretically by Castro and Trillas and developed into a training algor…
- Analog-Friendly Predictive Coding without Activation Derivatives
Francesco Innocenti · 29 septembre 2026
Predictive coding (PC) is a local, energy-based alternative to backpropagation (BP) whose iterative inference dynamics make it attractive for implementation on analog hardware. However, standard nonlinear PC requires evaluating the derivative of the activation function during both inference and lear…
- Emergent One-Third Scaling Law as Attention Tries to Concentrate
Yizhou Liu, Sara Kangaslahti, Jeff Gore · 29 septembre 2026
The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distr…
- LocalProp: Neuro-Localized Memory-Efficient Backpropagation
Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu · 29 septembre 2026
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly loca…
- Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
Jaehee Seo, Jisu Kim · 28 septembre 2026
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existin…
- Common-Mode Collapse and Recovery in Direct Feedback Alignment
Varun Reddy, Bernardo L. Sabatini, Houman Safaai · 28 septembre 2026
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's co…
- The Residual Stream's Effective Depth
Barak Gahtan, Ido Galil, Alex M. Bronstein · 28 septembre 2026
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only langu…
- LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe · 28 septembre 2026
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each la…
- Population loss in shallow ReLU networks: Bias & families of critical points
Michael Field · 28 septembre 2026
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The ne…
- Staged Depth Training: A Representation Curriculum for PINNs
Kejia Zhang, Youran Sun, Haizhao Yang · 28 septembre 2026
Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representations are explicitly learned, transferred …
- NeuralCert: certified computational discovery of extremal mathematical constructions
Mark Patrick Roeling · 28 septembre 2026
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable repr…
Autres sujets du thème Intelligence artificielle
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Large Language Models7 407 papiers / 12 mois+247 %
- Adversarial Robustness in Machine Learning3 552 papiers / 12 mois+118 %
- Reinforcement Learning in Robotics2 519 papiers / 12 mois+117 %
- Explainable Artificial Intelligence (XAI)2 319 papiers / 12 mois+200 %
- Domain Adaptation and Few-Shot Learning2 059 papiers / 12 mois+67 %
- Advanced Graph Neural Networks1 926 papiers / 12 mois+38 %
