Physical Sciences › Computer Science › Artificial Intelligence
Neural Networks and Applications
245 papers indexed
Artificial neural networks explore learning and optimization mechanisms through various approaches, such as enhancing data with techniques like Stochastic Weight Averaging or studying minimal models to explain phenomena like scaling laws. Some works focus on the very structure of these networks, analyzing their curvature, symmetry, or activation functions, while others examine specific architectures like Probabilistic Circuits or dual-encoder heads to reduce complexity. Still others address theoretical or practical questions, such as performance certification, adaptation to novel object orientations, or the integration of physical constraints into learning.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States28% · 18 papers
- China14% · 9 papers
- France9.2% · 6 papers
- Germany9.2% · 6 papers
- United Kingdom7.7% · 5 papers
- Spain7.7% · 5 papers
- Sweden6.2% · 4 papers
- Canada6.2% · 4 papers
Across 65 papers on this subject with at least one lab located. 28 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning
Syon Mansur, Joel Zylberberg · 2 October 2026
End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes ach…
- Misalignment of Low-Loss Regions Causes Grokking
Yongding Tian, Zaid Al-Ars, Maksim Kitsak, Peter Hofstee · 2 October 2026
Grokking refers to the delayed emergence of validation-set generalization after a model has already overfit the training set. Although first observed in small algorithmic tasks trained with transformers, its underlying mechanism remains unsettled. In this work, we develop an analysis framework based…
- Beyond Affine Transformations: A Soft Dominance Layer for Coordinate-Wise Neural Computation
Mariano Rivera · 2 October 2026
This paper presents a preliminary study of an alternative to the affine transformation underlying conventional neural-network layers. In the proposed Soft Dominance Layer, each output unit compares input coordinates with a learnable reference vector and aggregates smooth inequality responses. A sigm…
- The Life Cycle of a Massive Activation: Stochastic Birth, Weight-Decay-Driven Growth, and Competitive Consolidation
S. Aaron McClendon, Jorge Gallego-Feliciano, Antonios Saravanos · 2 October 2026
Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled intervent…
- Per-Node Activation Function Evolution in Indirectly Encoded Substrates: Solvability, Limits, and Emergent Diversity
Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel · 2 October 2026
Biological neurons achieve computational diversity through specialized types: tonic, bursting, adapting, and fast-spiking cells coexist within the same circuit. Artificial neural networks, by contrast, apply a single activation function uniformly to all nodes, which limits what they can represent. W…
- SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials
Amirhosein Azarpour, Seyyed Moein Kazemi · 2 October 2026
Kolmogorov-Arnold Networks (KANs) represent a paradigmatic shift in deep learning by replacing fixed node activations with learnable univariate functions on edges, offering enhanced interpretability and parameter efficiency. While recent polynomial-based KAN variants have addressed the computational…
- The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State · 1 October 2026
Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states f…
- Patch Rebirth: Fast and Transferable Model Inversion of Vision Transformers
Seongsoo Heo, Dong-Wan Choi · 1 October 2026
Model inversion is a widely adopted technique in data-free learning that reconstructs synthetic inputs from a pretrained model through iterative optimization, without access to original training data. Unfortunately, its application to state-of-the-art Vision Transformers (ViTs) poses a major computa…
- Effective Does Not Mean Useful: Conditional Functional Substitutability for Redundancy and Scaling in Transformers
Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Xinyong Cai, Juncheng Bu, Lan Yu, Tinghe Zhang · 1 October 2026
Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We view redundancy as an input-conditioned, dynamic relation: intermediat…
- How Local Mixing Encodes Relative Position in Global NoPE Attention
Cutter Dawes, Nick Alonso, Tom Figliolia, Beren Millidge · 30 September 2026
The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encod…
- Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
Zahra Khodagholi, Niloofar Yousefi · 30 September 2026
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting t…
- Task-Relevant Null-Space Residuals for Non-Injective Neural Mappings
Bizu Feng, Zhimu Yang, Shuming Wang, Yuan Cheng, Shaode Yu, Xiaojun Qian, Zixin Hu · 30 September 2026
Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-…
- Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
Soutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron, Nadav Dym · 29 September 2026
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical founda…
- From Distributions to Stochastic Processes: Neural Approximation of Measure-Valued Maps
Yichen Wang, Ziyi Wang, Wenlian Lu, Chenghuang Shen, Jianfeng Liu, Zhengdong Xiao, Longjiu Luo, Qianrong Wang · 29 September 2026
Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learning and extend it to mappings between stoc…
- \L{}ukasiewicz Neural Networks Extended: Residual Architectures and Crystallization Strategies for Interpretable Rule Extraction
Carlos Leandro · 29 September 2026
A feed-forward neural network whose weights are integers and whose activation is the truncated identity implements, neuron by neuron, the connectives of \L{}ukasiewicz many-valued logic. This exact correspondence --- established theoretically by Castro and Trillas and developed into a training algor…
- Analog-Friendly Predictive Coding without Activation Derivatives
Francesco Innocenti · 29 September 2026
Predictive coding (PC) is a local, energy-based alternative to backpropagation (BP) whose iterative inference dynamics make it attractive for implementation on analog hardware. However, standard nonlinear PC requires evaluating the derivative of the activation function during both inference and lear…
- Emergent One-Third Scaling Law as Attention Tries to Concentrate
Yizhou Liu, Sara Kangaslahti, Jeff Gore · 29 September 2026
The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distr…
- LocalProp: Neuro-Localized Memory-Efficient Backpropagation
Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu · 29 September 2026
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly loca…
- Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
Jaehee Seo, Jisu Kim · 28 September 2026
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existin…
- Common-Mode Collapse and Recovery in Direct Feedback Alignment
Varun Reddy, Bernardo L. Sabatini, Houman Safaai · 28 September 2026
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's co…
- The Residual Stream's Effective Depth
Barak Gahtan, Ido Galil, Alex M. Bronstein · 28 September 2026
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only langu…
- LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe · 28 September 2026
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each la…
- Population loss in shallow ReLU networks: Bias & families of critical points
Michael Field · 28 September 2026
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The ne…
- Staged Depth Training: A Representation Curriculum for PINNs
Kejia Zhang, Youran Sun, Haizhao Yang · 28 September 2026
Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representations are explicitly learned, transferred …
- NeuralCert: certified computational discovery of extremal mathematical constructions
Mark Patrick Roeling · 28 September 2026
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable repr…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
