Physical Sciences › Computer Science › Artificial Intelligence
Neural Networks and Applications
245 artículos indexados
Las redes neuronales artificiales exploran mecanismos de aprendizaje y optimización a través de enfoques variados, como la mejora de datos mediante técnicas tales como el Stochastic Weight Averaging o el estudio de modelos mínimos para explicar fenómenos como las leyes de escala. Algunos trabajos se centran en la estructura misma de estas redes, analizando su curvatura, su simetría o sus funciones de activación, mientras que otros examinan arquitecturas específicas como los Probabilistic Circuits o los dual-encoder heads para reducir su complejidad. Otros abordan cuestiones teóricas o prácticas, como la certificación del rendimiento, la adaptación a orientaciones inéditas de objetos o la integración de restricciones físicas en el aprendizaje.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos28 % · 18 artículos
- China14 % · 9 artículos
- Francia9,2 % · 6 artículos
- Alemania9,2 % · 6 artículos
- Reino Unido7,7 % · 5 artículos
- España7,7 % · 5 artículos
- Suecia6,2 % · 4 artículos
- Canadá6,2 % · 4 artículos
Sobre 65 artículos de este tema con al menos un laboratorio localizado. 28 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State · 1 de octubre de 2026
Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states f…
- Patch Rebirth: Fast and Transferable Model Inversion of Vision Transformers
Seongsoo Heo, Dong-Wan Choi · 1 de octubre de 2026
Model inversion is a widely adopted technique in data-free learning that reconstructs synthetic inputs from a pretrained model through iterative optimization, without access to original training data. Unfortunately, its application to state-of-the-art Vision Transformers (ViTs) poses a major computa…
- Effective Does Not Mean Useful: Conditional Functional Substitutability for Redundancy and Scaling in Transformers
Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Xinyong Cai, Juncheng Bu, Lan Yu, Tinghe Zhang · 1 de octubre de 2026
Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We view redundancy as an input-conditioned, dynamic relation: intermediat…
- How Local Mixing Encodes Relative Position in Global NoPE Attention
Cutter Dawes, Nick Alonso, Tom Figliolia, Beren Millidge · 30 de septiembre de 2026
The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encod…
- Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
Zahra Khodagholi, Niloofar Yousefi · 30 de septiembre de 2026
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting t…
- Task-Relevant Null-Space Residuals for Non-Injective Neural Mappings
Bizu Feng, Zhimu Yang, Shuming Wang, Yuan Cheng, Shaode Yu, Xiaojun Qian, Zixin Hu · 30 de septiembre de 2026
Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-…
- Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
Soutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron, Nadav Dym · 29 de septiembre de 2026
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical founda…
- From Distributions to Stochastic Processes: Neural Approximation of Measure-Valued Maps
Yichen Wang, Ziyi Wang, Wenlian Lu, Chenghuang Shen, Jianfeng Liu, Zhengdong Xiao, Longjiu Luo, Qianrong Wang · 29 de septiembre de 2026
Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learning and extend it to mappings between stoc…
- \L{}ukasiewicz Neural Networks Extended: Residual Architectures and Crystallization Strategies for Interpretable Rule Extraction
Carlos Leandro · 29 de septiembre de 2026
A feed-forward neural network whose weights are integers and whose activation is the truncated identity implements, neuron by neuron, the connectives of \L{}ukasiewicz many-valued logic. This exact correspondence --- established theoretically by Castro and Trillas and developed into a training algor…
- Analog-Friendly Predictive Coding without Activation Derivatives
Francesco Innocenti · 29 de septiembre de 2026
Predictive coding (PC) is a local, energy-based alternative to backpropagation (BP) whose iterative inference dynamics make it attractive for implementation on analog hardware. However, standard nonlinear PC requires evaluating the derivative of the activation function during both inference and lear…
- Emergent One-Third Scaling Law as Attention Tries to Concentrate
Yizhou Liu, Sara Kangaslahti, Jeff Gore · 29 de septiembre de 2026
The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distr…
- LocalProp: Neuro-Localized Memory-Efficient Backpropagation
Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu · 29 de septiembre de 2026
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly loca…
- Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
Jaehee Seo, Jisu Kim · 28 de septiembre de 2026
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existin…
- Common-Mode Collapse and Recovery in Direct Feedback Alignment
Varun Reddy, Bernardo L. Sabatini, Houman Safaai · 28 de septiembre de 2026
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's co…
- The Residual Stream's Effective Depth
Barak Gahtan, Ido Galil, Alex M. Bronstein · 28 de septiembre de 2026
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only langu…
- LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe · 28 de septiembre de 2026
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each la…
- Population loss in shallow ReLU networks: Bias & families of critical points
Michael Field · 28 de septiembre de 2026
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The ne…
- Staged Depth Training: A Representation Curriculum for PINNs
Kejia Zhang, Youran Sun, Haizhao Yang · 28 de septiembre de 2026
Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representations are explicitly learned, transferred …
- NeuralCert: certified computational discovery of extremal mathematical constructions
Mark Patrick Roeling · 28 de septiembre de 2026
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable repr…
- Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation
Venkata Subbaiah Yerrapati, Rahul Dixit, Ajay Kumar Shukla · 28 de septiembre de 2026
Understanding the features captured by the hidden layers of neural networks is a fundamental challenge in machine learning, despite their widespread success across various classification problems. In this work, we propose an algebraic framework for examining neural networks that model classification…
- Orbital Error Dynamics: Self-Organized Criticality, Ephemeral Parameter Resonance, and Non-Linear Biological Ontologies in Zero-Storage Neural Synthesis
Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i} · 25 de septiembre de 2026
Modern deep neural networks treat parameters as static floating-point matrices stored in physical memory, incurring Von Neumann memory bottlenecks and representation collapse. We formulate Orbital Error Dynamics (OED), an analytical framework wherein synaptic weights are not stored masses (O(W)), bu…
- ELiSe: Efficient Learning of Sequences in Structured Recurrent Networks
Laura Kriener, Kristin V\"olk, Ben von H\"unerbein, Federico Benitez, Walter Senn, Mihai A. Petrovici · 25 de septiembre de 2026
Behavior can be described as a temporal sequence of actions driven by neural activity. To learn complex sequential patterns in neural networks, memories of past activities need to persist on significantly longer timescales than the relaxation times of single-neuron activity. While recurrent networks…
- Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking
Zhiyu Zhang, Yupeng Li · 25 de septiembre de 2026
State tracking requires composing a sequence of updates, but accuracy alone does not reveal what a model has learned. We study neural networks trained to predict the running product of group elements. We identify quotient solutions in Transformers, where models recover the quotient class while predi…
- Statistical Properties of Deep Neural Networks with Dependent Data
Chad Brown · 24 de septiembre de 2026
This paper develops theory for deep neural network (DNN) estimators under dependent data. To provide theory applicable to a variety of DNN-based estimators, I first establish nonasymptotic probability bounds on the theoretical and empirical $\mathcal{L}^{2}$-errors of nonparametric sieve estimators …
- Binary Quantized Neural Network Training Is W[1]-Hard Parameterized by Input and Output Dimensions
Tao Jiang, Minbo Gao, Shaowei Cai · 24 de septiembre de 2026
Ganian et al. (ICLR 2026) proved that quantized neural network training is fixed-parameter tractable when parameterized jointly by architecture treewidth, input dimension $\alpha$, and output dimension $\omega$, and left open whether $\alpha+\omega$ alone yields fixed-parameter tractability. We prov…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
