AI Research
arXiv paper trends
Recent submissions across eight fields, including AI, machine learning, space and energy
Category distribution
- Machine Learning(13)
- Computer Vision(12)
- Quantum Physics(6)
- Robotics(4)
- cs.CL(2)
- AI(2)
- Energy Systems(1)
Papers
- Computer VisionMoore, Escher, Penrose: A Conformal Golden Braid
I don't think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can th
- Computer VisionSphere Encoder 2
Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generati
- Computer VisionOne Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained
- cs.CLKaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessment
- Computer VisionROWBench: Do Video Models Render What the Program Specifies?
Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and i
- RoboticsReconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruc
- Computer VisionEmbedding Prediction Helps Image Generation
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition in
- AIScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of res
- Computer VisionSILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragme
- AIVISTA: A Visual Harness for Reasoning in an Interactive World
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA,
- Machine LearningTACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress opt
- Machine LearningFERPO: Forward Entropy-Regularized Policy Optimization
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict
- Computer VisionHiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation
Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which a
- RoboticsInterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining wh
- Machine LearningCost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
The generalized Schrödinger bridge on a graph moves mass between two distributions while charging a cost for the states visited. It has been approached by learning the rates of a controlled continuous
- cs.CLHierarchical Continuous Diffusion Language Models
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structur
- Machine LearningThe Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlyi
- Machine LearningTrust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and
- Machine LearningGenerative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional des
- Computer VisionDMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to th
- Energy SystemsMinimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
On broad classes of linear systems, the shortest experiments are almost as good as the best possible ones. For $n$ states and $m$ inputs, the shortest input sequences that support robust data-driven s
- Machine LearningHigher-Order Molecular Grammars for Generative and Foundation Models in Chemistry
Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring system
- Machine LearningDecoding Looped Transformers Better for (Almost) Free
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet s
- Quantum PhysicsSingle-Particle Spectral Estimation
We study Hamiltonian learning in a novel setting, where we are promised a physical structure that is obfuscated by a global unitary. In particular, we consider the task of learning the weights $E_i$ o
- Machine LearningSoftServe: A Scalable Quasi-Newton Method for Deep Learning
Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and en
- Computer VisionOmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an e
- Computer VisionGenerative Cinematographer: Composing Camera and Object Motion in 3D
Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond
- Machine LearningFrom Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Q
- Quantum PhysicsLearning SYK Hamiltonians
We study the problem of learning the dense Sachdev--Ye--Kitaev (SYK) Hamiltonian from copies of its Gibbs state. Existing algorithms for Hamiltonian learning typically rely on geometric locality or bo
- Quantum PhysicsFrom Permutation Symmetry to Communication Bounds and Additivity
Correlations across channel uses can improve quantum communication rates, making optimization over arbitrarily large blocks a central difficulty in determining quantum capacity. We show how full permu
- Machine LearningEffective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes
Protein function annotation needs to know which predictions to distrust, not only what a model predicts. We ask whether tissue-specific interaction structure carries that information. Our candidate si
- Machine LearningEvery Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear
- Quantum PhysicsRobust exponential lower bounds for fermionic and bosonic Gaussian ranks
The power and limitations of classical simulation are central to understanding quantum computational advantages. A leading simulation paradigm is based on coherent decomposition into classically tract
- RoboticsWatch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination
Robots operating in the physical world will increasingly need to coordinate with other robots, particularly in manipulation tasks where an object may be too large or heavy for a single robot to carry
- Quantum PhysicsPolynomial-time classical and quantum simulation of quantum impurity models
Quantum impurity models are paradigmatic models of interacting quantum matter, as well as key computational primitives for modern electronic-structure methods. They describe a small subsystem of inter
- Quantum PhysicsBeyond Light Cones: State Preparation Complexity in Quantum Spin Glasses
We introduce a method for studying state preparation complexity in dense quantum $p$-spin Hamiltonians on $n$ qubits, going beyond bounds based only on circuit lightcones. The key input is the class's
- Computer VisionWorld Observer: Joint Actor-Observer Generation for Persistent World Modeling
How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agent's actions, yet remain actor-centric. Once an ob
- RoboticsDuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending
- Computer Vision4Director: Controlling Video World Models with Rigid 3D Geometry
Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and r
- Machine LearningWhen Do Intrinsic Rewards Lead to Exploration?
Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizin