Physical Sciences › Computer Science › Artificial Intelligence
Neural Networks and Applications
245 indexierte Paper
Künstliche neuronale Netze erforschen Lern- und Optimierungsmechanismen durch verschiedene Ansätze, wie die Verbesserung von Daten durch Techniken wie Stochastic Weight Averaging oder die Untersuchung minimaler Modelle zur Erklärung von Phänomenen wie Skalierungsgesetzen. Einige Arbeiten konzentrieren sich auf die Struktur dieser Netze selbst, indem sie deren Krümmung, Symmetrie oder Aktivierungsfunktionen analysieren, während andere spezifische Architekturen wie Probabilistic Circuits oder dual-encoder heads untersuchen, um deren Komplexität zu reduzieren. Wieder andere behandeln theoretische oder praktische Fragen, wie die Zertifizierung von Leistungen, die Anpassung an ungewöhnliche Objektausrichtungen oder die Integration physikalischer Constraints in das Lernen.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten28 % · 18 Artikel
- China14 % · 9 Artikel
- Frankreich9,2 % · 6 Artikel
- Deutschland9,2 % · 6 Artikel
- Vereinigtes Königreich7,7 % · 5 Artikel
- Spanien7,7 % · 5 Artikel
- Schweden6,2 % · 4 Artikel
- Kanada6,2 % · 4 Artikel
Über 65 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 28 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning
Syon Mansur, Joel Zylberberg · 2. Oktober 2026
End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes ach…
- Misalignment of Low-Loss Regions Causes Grokking
Yongding Tian, Zaid Al-Ars, Maksim Kitsak, Peter Hofstee · 2. Oktober 2026
Grokking refers to the delayed emergence of validation-set generalization after a model has already overfit the training set. Although first observed in small algorithmic tasks trained with transformers, its underlying mechanism remains unsettled. In this work, we develop an analysis framework based…
- Beyond Affine Transformations: A Soft Dominance Layer for Coordinate-Wise Neural Computation
Mariano Rivera · 2. Oktober 2026
This paper presents a preliminary study of an alternative to the affine transformation underlying conventional neural-network layers. In the proposed Soft Dominance Layer, each output unit compares input coordinates with a learnable reference vector and aggregates smooth inequality responses. A sigm…
- The Life Cycle of a Massive Activation: Stochastic Birth, Weight-Decay-Driven Growth, and Competitive Consolidation
S. Aaron McClendon, Jorge Gallego-Feliciano, Antonios Saravanos · 2. Oktober 2026
Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled intervent…
- Per-Node Activation Function Evolution in Indirectly Encoded Substrates: Solvability, Limits, and Emergent Diversity
Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel · 2. Oktober 2026
Biological neurons achieve computational diversity through specialized types: tonic, bursting, adapting, and fast-spiking cells coexist within the same circuit. Artificial neural networks, by contrast, apply a single activation function uniformly to all nodes, which limits what they can represent. W…
- SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials
Amirhosein Azarpour, Seyyed Moein Kazemi · 2. Oktober 2026
Kolmogorov-Arnold Networks (KANs) represent a paradigmatic shift in deep learning by replacing fixed node activations with learnable univariate functions on edges, offering enhanced interpretability and parameter efficiency. While recent polynomial-based KAN variants have addressed the computational…
- The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State · 1. Oktober 2026
Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states f…
- Patch Rebirth: Fast and Transferable Model Inversion of Vision Transformers
Seongsoo Heo, Dong-Wan Choi · 1. Oktober 2026
Model inversion is a widely adopted technique in data-free learning that reconstructs synthetic inputs from a pretrained model through iterative optimization, without access to original training data. Unfortunately, its application to state-of-the-art Vision Transformers (ViTs) poses a major computa…
- Effective Does Not Mean Useful: Conditional Functional Substitutability for Redundancy and Scaling in Transformers
Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Xinyong Cai, Juncheng Bu, Lan Yu, Tinghe Zhang · 1. Oktober 2026
Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We view redundancy as an input-conditioned, dynamic relation: intermediat…
- How Local Mixing Encodes Relative Position in Global NoPE Attention
Cutter Dawes, Nick Alonso, Tom Figliolia, Beren Millidge · 30. September 2026
The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encod…
- Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
Zahra Khodagholi, Niloofar Yousefi · 30. September 2026
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting t…
- Task-Relevant Null-Space Residuals for Non-Injective Neural Mappings
Bizu Feng, Zhimu Yang, Shuming Wang, Yuan Cheng, Shaode Yu, Xiaojun Qian, Zixin Hu · 30. September 2026
Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-…
- Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
Soutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron, Nadav Dym · 29. September 2026
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical founda…
- From Distributions to Stochastic Processes: Neural Approximation of Measure-Valued Maps
Yichen Wang, Ziyi Wang, Wenlian Lu, Chenghuang Shen, Jianfeng Liu, Zhengdong Xiao, Longjiu Luo, Qianrong Wang · 29. September 2026
Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learning and extend it to mappings between stoc…
- \L{}ukasiewicz Neural Networks Extended: Residual Architectures and Crystallization Strategies for Interpretable Rule Extraction
Carlos Leandro · 29. September 2026
A feed-forward neural network whose weights are integers and whose activation is the truncated identity implements, neuron by neuron, the connectives of \L{}ukasiewicz many-valued logic. This exact correspondence --- established theoretically by Castro and Trillas and developed into a training algor…
- Analog-Friendly Predictive Coding without Activation Derivatives
Francesco Innocenti · 29. September 2026
Predictive coding (PC) is a local, energy-based alternative to backpropagation (BP) whose iterative inference dynamics make it attractive for implementation on analog hardware. However, standard nonlinear PC requires evaluating the derivative of the activation function during both inference and lear…
- Emergent One-Third Scaling Law as Attention Tries to Concentrate
Yizhou Liu, Sara Kangaslahti, Jeff Gore · 29. September 2026
The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distr…
- LocalProp: Neuro-Localized Memory-Efficient Backpropagation
Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu · 29. September 2026
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly loca…
- Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
Jaehee Seo, Jisu Kim · 28. September 2026
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existin…
- Common-Mode Collapse and Recovery in Direct Feedback Alignment
Varun Reddy, Bernardo L. Sabatini, Houman Safaai · 28. September 2026
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's co…
- The Residual Stream's Effective Depth
Barak Gahtan, Ido Galil, Alex M. Bronstein · 28. September 2026
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only langu…
- LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe · 28. September 2026
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each la…
- Population loss in shallow ReLU networks: Bias & families of critical points
Michael Field · 28. September 2026
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The ne…
- Staged Depth Training: A Representation Curriculum for PINNs
Kejia Zhang, Youran Sun, Haizhao Yang · 28. September 2026
Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representations are explicitly learned, transferred …
- NeuralCert: certified computational discovery of extremal mathematical constructions
Mark Patrick Roeling · 28. September 2026
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable repr…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
