BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
Reward-Density Heuristic for Dynamic Multi-Vehicle Routing: Performance and Computational Efficiency
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
Scalable Perturbation Learning for Online Self-Supervised Echo State Networks
Modeling Normal Is All You Need: Joint Latent Clustering for Anomaly Detection in Multimodal Cyber-Physical Systems
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development
EcoVision: AI-Powered Drone Imaging for Salt Marsh Vegetation Monitoring and Dominance Mapping
LLM-Guided Measurement Credibility Correction for Trustworthy Industrial Process Inference
X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
A toy framework for single and multi-agent human-AI curiosity ecosystems
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Canopy: A Heterograph Foundation Model for Metabolic Engineering
Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval
From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition
Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail
Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models
Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction
TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting
Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers
Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement
A Function-Space Dichotomy for Compositional Learning: Exponential Sub-Optimality of the Neural Tangent Kernel
What Images Cannot Say: Language-Guided Olfactory Representation Learning
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders
From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
Provable learning separation for predicting time-evolution of quantum many-body systems
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment
Graph Convolutional Attention: A Spectral Perspective on Graph Denoising and Diffusion
Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses
Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)
RUFNet: Query-Guided Support Mask Refinement and Uncertainty Fusion based on Hybrid Mamba for Few-Shot Brain Tumor Segmentation
CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion
Toward Trustworthy Large Language Model Agents in Healthcare
Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization
Computing Monetary Risk Measures in Linear Time
AIFS-SUBS: Extending Data-Driven Forecasting to Sub-Seasonal Timescales
Choosing a parallel heterogeneous ensemble method for tabular classification
Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters
Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing
PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization
Latent Programming Horizons in Coding Agents
Noisy-Channel Minimum Bayes Risk Decoding
EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
Probing Geospatial SSL Representations with Environmental Signals
Video-based detection of cessation of breathing in pre-term infants using machine learning
MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models
Streaming Neural Speech Codecs through Time-Invariant Representations
FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation
SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment Analysis
Target-Guided Selective Reweighting for Physics-Informed Neural Network Inverse Problems: A Transfer Learning Approach
Untrusted Content Masking for Web Agents with Security Guarantees
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
Learning Only What Valid Adapters Can Express: Subspace-Constrained Adaptation Against Fine-Tuning Poisoning
Biologically Informed Deep Neural Networks for Multi-Omic Integration, Pathway Activity Inference and Risk Stratification in Cancer
Quantum Spectral Anomaly Detection
Evaluating and Understanding Model Editing for Medical Vision Language Models
How Much is Left? LLMs Linearly Encode Their Remaining Output Length
How Far is Too Far? Defining the Distance Threshold for Verification Siamese Networks
OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
TabPack: Efficient Hyperparameter Ensembles for Tabular Deep Learning
What Does a Discrete Diffusion Model Learn?
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling
AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes
Environmental Drivers of Respiratory Disease: A District Level Analysis
Transferability Between Understanding and Generation in Unified Multimodal Models
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
dOPSD: On-Policy Self-Distillation for Diffusion Language Models
evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations
Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models
Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models
Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning
Two Black Boxes, One Solver: Encoder Probing and Decoder Attribution for Neural Multi-Attribute VRP under Hard-Mask and Recourse Decoders
From Interaction to Intent: Inferring User Objectives from Provenance Logs
VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models
Beyond travel mode: urban context shapes active mobility's mental health effects over time
Training-Free Model Selection and Domain-Aware Score Calibration for First-Shot Anomalous Sound Detection
Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents
Mechanism-level routing failure in LLMs over Lean-verified algebraic structures
CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining
Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
Explainable Novel Category Discovery in Semantic Concept Space
MTEB-PT: A Text Embedding Benchmark for Brazilian Portuguese
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection
G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement
SILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable Routing
Machine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based Methodology
GlaKG: A Biomarker-Centric Fundus Knowledge Graph for Explainable Glaucoma Diagnosis and Risk Assessment
URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment
PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection
yetone/native-feel-skill:An Agent Skill for designing cross-platform desktop apps that feel native — distilled from Raycast's 2.0 deep-dive and reverse engineering of Raycast Beta.app. Eight architectural tenets, four-layer architecture, WebKit/WebView2 survival guide, 75-item ship audit.
rpamis/comet:Comet: agent skill harness phase-guarded automation from idea to archive
Show HN: TaskPeace – a task queue my AI coding agents pull work from over MCP
Coding-agents can replicate scientific machine learning papers
Probing Chemical Language Models: Effects of Pre-training and Fine-tuning
Dynamic Neural Graph Encoding of Inference Processes in Deep Weight Space
Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation
RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation
UA-ChatDev: Uncertainty-Aware Multi-Agent Collaboration for Reliable Software Development
An Optimisation Framework for the Well-Conditioned Training of Physics-Informed Neural Networks
Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond
The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
Prediction Sets for Counterfactual Decisions: Coverage, Optimality, and Conformal Prediction
Criticality-Based Guard Rail Validation for AI Agent Decisions in Autonomous Telecom Networks
An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility
Aggregation with Exponential Weights is Optimal in Expectation
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures
Dendritic In-Context Learning in a Single-Layer Spiking Neural Network
Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces
GAP-GDRNet: Geometry-Aware Monocular Visual Pose Sensing on a Single-Target Synthetic Spacecraft Dataset
Stable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory Gates
VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval
Know Your Source: A Public Knowledge Store for Media Background Checks
Bringing Agentic Search to Earth Observation Data Discovery
Steerability via constraints: a substrate for scalable oversight of coding agents
ACID: Action Consistency via Inverse Dynamics for Planning with World Models
Q-GAIN: A Python Package for Machine Learning and Physically Informed Analysis Applications
The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing
Neuron-Aware Active Few-Shot Learning for LLMs
WorldSample: Closed-loop Real-robot RL with World Modelling
Extreme Adaptive Transformer for Time Series Forecasting
Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID Data
Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning
Controllable Sim Agents with Behavior Latents
Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas
What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates
Online Safety Monitoring for LLMs
LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
Show HN: QUALITY.md – open format/specification, agent skill, and CLI
Show HN: Mail Memories – A desktop app to rescue photos from Gmail
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models
Dynamic Bidirectional Pattern Memory: A Production-Scale Empirical Characterisation of Inference-Time Gating in Clinical NLP
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm
Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions
Aionoscope: Debugging Latent-State Accessibility in Time-Series Representations
Bridging Quantum Computing Paradigms toward Semiconductor Yield: A Controlled CV-versus-DV Comparison on Wafer-Map Defect Classification
TRCGL-Net: A Long-Tailed Multi-Label Chest X-Ray Classification Framework with Generative Data Augmentation and Label Co-Occurrence Modeling
Foundation Models vs. Radiomics for Lung Computed Tomography: A Benchmark of Feature Extractors, Classification Heads, and Segmentation Choices
Understanding Large Language Models
Evidence-Supported Credit Risk Report Generation Using News-Centric Financial Knowledge Graphs
Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework
EchoRisk: A Multicentre Echocardiography Dataset and Benchmark for Cardio-Oncology
Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Message Passing Enables Efficient Reasoning
Balancing Expressivity and Learnability in Quantum Kernel Bandit Optimization
Group-invariant Coresets for Data-efficient Active Learning
SynLaD: Latent Diffusion for Generating Synthesizable Molecules Conditioned on 3D Pharmacophore Profiles
FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement
GAIA: Geometry-Adaptive Operator Learning for Forward and Inverse Problems
Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search
A Lightweight Self-Supervised Learning Framework for Multivariate Time Series using Hierarchical-JEPA on ECG Data
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
Decision-Aware Training for Sample-Based Generative Models
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations
Optimal Resource Utilization for Autonomous Laboratory Orchestrators
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model
The State-Prediction Separation Hypothesis
Language-Critique Imitation Learning from Suboptimal Demonstrations
Measuring the Gap Between Human and LLM Research Ideas
Show HN: AnalystAIPack – 118 runnable agent skills for malware analysis and RE
cloudflare/security-audit-skill:A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection
Overview of the TalentCLEF 2026: Skill and Job Title Intelligence for Human Capital Management
ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping
Diffusing Blame: Task-Dependent Credit Assignment in Biologically Plausible Dual-Stream Networks
Nonlinearity-Aware LoRA: Structured Gate Adaptation under Low-Rank Constraints
Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian
STEB: Style Text Embedding Benchmark
JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering
A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols
Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
SpikeLogBERT: Energy-Efficient Log Parsing Using Spiking Transformer Networks
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision
Relational and Sequential Conformal Inference for Energy Time Series over Graphs via Foundation Models
Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models
Low-dimensional topology of deep neural networks
Accelerating Conformal Prediction via Approximate Leave-One-Out
Semantic Leakage and Privacy Preservation in Relay-Assisted Semantic Communications
TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization
Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA
CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation
Scalable Behaviour Cloning on Browser Using via Skill Distillation
FedLAB: Traceable Semantic Codebooks for Federated Multimodal Graph Foundation Learning
Automated Background Swapping for Robustness against Spurious Backgrounds
SemRF: A Semantic Reference Frame for Residual-Stream Dynamics in Language Models
Generative Skill Composition for LLM Agents
AdaJEPA: An Adaptive Latent World Model
Freeform Preference Learning for Robotic Manipulation
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision
Spatial Reasoning via Modality Switching Between Language and Symbolic Representation
From Idea to Prototype in an Afternoon: Scaffolded, AI-Assisted Rapid VA Prototyping
Safe Online Learning via Smooth Safety-Structured Policy Composition
PGUDA: Pressure-Guided Unsupervised Domain Adaptation with Cross-Modal Knowledge Distillation for sEMG-Based Gesture Recognition
Dualformer: Efficient Feature Extractor for Complex-valued Blind Communication Signal Analysis
Stage-Transition Dense Reward Modeling for Reinforcement Learning
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
World-Model Collapse as a Phase Transition
Xiaomi-GUI-0 Technical Report
DA-Studio: An Agentic System for End-to-End Data Analysis
Ask the World Before Acting: Budgeted Environment Probing for World-Model Calibration
CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models
CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market
Team MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline Dynamics
Von Mises Based Uncertainty Quantification for Closely Spaced Automotive Radar Targets
Surprise as a Signal for Plasticity and Metacognition
Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models
Design and Implementation of Agentic Orchestrations and Orchestration of Agents
RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference
On the Convergence of Self-Improving Online LLM Alignment
Beyond the Expressivity-Trainability Paradox: A Dynamical Lie Algebra Perspective on Navigating Barren Plateaus in Quantum Machine Learning
ACE: Pluggable Adaptive Context Elasticizer across Agents
Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning
DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers
Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning
Preserve the Hard, Regenerate the Rest: Uncertainty-Guided Synthetic Training Data Augmentation with Diffusion Models
A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems
FARS: A Fully Automated Research System Deployed at Scale
Extrapolating from Regularised Solutions for Solving Ill-Conditioned Linear Systems in Machine Learning
BayesEvolve: Explicit Belief States for Autonomous Scientific Discovery
FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks
REAR: Test-time Preference Realignment through Reward Decomposition
Early Cue Precision Shapes Visual Shortcut Learning in Controlled Cue-Manipulation Benchmarks
Residual-Guided Expert Specialization for Incomplete Multimodal Learning
OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL
On the Vulnerability of Parameter-Level Defenses to Model Merging
Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data
Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation
Scalar Representations of Neural Network Training Dynamics
Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding
Beyond IID: How General Are Tabular Foundation Models, Really?
Can LLMs Rank? A Tale of Triads and Triage
Beyond Point Estimates for Glaucoma Visual Field Forecasting with Diffusion Models
Translating Natural Language to Strategic Temporal Specifications via LLMs
The FIL Hypothesis: Inductive Biases Help with Kernel Engineering
On the Faithfulness of Post-Hoc Concept Bottleneck Models
Discovering Collaboration from Novelty: Random Network Distillation for Clustered Federated Learning
$μ$Flow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors
Entity Binding Failures in Tool-Augmented Agents
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?
Convergence of Continual Learning in Homogeneous Deep Networks
The Human Creativity Benchmark
A Multi-task Mixture of Experts Framework for Malware Classification, Packing Detection, and Family Attribution
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
The Fundamental Limits of Valid Transport Map Estimation
Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization
Uncertainty-Aware Generation and Decision-Making Under Ambiguity
Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection
Wireless Backdoor Attack and Defense for Semantic Communications over Multiple Access Channel
Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms
DOPD: Dual On-policy Distillation
GROW$^2$: Grounding Which and Where for Robot Tool Use
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
Exploiting Local Flatness for Efficient Out-of-Distribution Detection
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
SWE-Together: Evaluating Coding Agents in Interactive User Sessions
DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation
RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines
NeuReasoner: Theory-grounded Mapping of Reasoning Elicitation Boundaries
First-Order Temporal Logic Tensor Networks
Exploration and Online Transfer with Behavioral Foundation Models
Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping
AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills
T3R: Deeper Test-Time Adaptation for Graph Neural Networks via Gradient Rotation
Consensus Clustering of Free-Viewing Gaze Data: New Insights into Human-Information Interaction
Heads, Not Backbones: Output Heads Dominate Architectures on Fat-Tailed Returns
Building Multi-Task Agentic LLMs via Two-Phase Distillation
Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance
Data-Driven Energy-Based Learning via Gibbs Measures on Hierarchical Structures
Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts
Efficient Retrieval-Augmented Generation via Token Co-occurrence Graphs
Temporal Feature Extractors in EEG Foundation Models: A Controlled Comparison Including a Pretrained Time-Series Model
Automating the Design of Embodied AgentArchitectures
SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance
Gravitational Duals from Equations of State II: Large Hierarchies and False Vacua
Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters
Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching
DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks
Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning
From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent
Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data
Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning
B3O: Scalable Boltzmann Batch Bayesian Optimization
A Distributionally Robust Framework for Learned Reconstructions in Inverse Problems
CaresAI at CT-DEB26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Models
Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
KnowsTFM: Knowledge-Informed Fine-Tuning of Small Tabular Foundation Models
Defending Against Harmful Supervision Hidden in Benign Samples
Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation
PromptGNN-sim: Deep Fusion and Alignment of GNN and LLMs for Text-Attributed Graph Learning
ManimAgent: Self-Evolving Multimodal Agents for Visual Education
Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
Forward-Future/loopy:A library of practical AI-agent loops and an installable skill for finding, adapting, and designing repeatable agent workflows.
Applicability of memorization indicators for early spotting of overfitting while recalibrating sEMG-decoders on low sample sizes
S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts
Mosaic: A Benchmark Suite for Differentiable Physics Solvers
Home3D 1.0: A High-Fidelity Image-to-3D Asset Generation System for Interior Design
Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy
AI Persuasive Framing in Collective Dilemmas
An Empirical Analysis of Factual Errors in Human-Written Text and its Application
Reasoning Beyond Prediction: From Data-Driven to Causal Software Engineering
ToxiREX: A Dataset on Toxic REasoning in ConteXt
Benchmarking on Tasks That Matter: Dataset Selection for Preserving Model Rankings
Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection
Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA
ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
Single and Multi Truth Data Fusion using Large Language Models
JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training
When One Adapter Speaks for Many: Discovering Low-Rank Redundancy in Continual Fine-Tuning
Beyond Sparse Supervision: Diffusion-Guided Learning for Few-Shot Graph Fraud Detection
Toward Robust In-Context Segmentation via Concept Guidance
Regularized Reward-Punishment Reinforcement Learning
EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
The Remittance Blueprint: Data-driven Intelligence for Sri Lanka
COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives
Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software
Bridging Ab Initio Symmetries and Global Nuclear Masses with Interpretable Neural Networks
Democratic ICAI: Debating Our Way to Steering Principles from Preferences
VGB for Masked Diffusion Model: Efficient Test-time Scaling for Reward Satisfaction and Sample Editing
Which Nash Equilibrium? Solver-Dependent Selection on Zero-Sum Nash Polytopes
DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand
AI boom risks global financial crash, warn central bankers
Task Failed Successfully: Saturating NIC and Disk Bandwidth
Aircraft crashes into Beijing's tallest skyscraper, triggering evacuations
XMSE-Aware Adaptive Empirical Bayes Estimation
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics
Decision-Aligned Evaluation of Uncertainty Quantification
Event-Aware Instructed Assistant for Referring Video Segmentation
On-board Remote-Sensing Foundation Models for Unsupervised Change Detection of Disaster Events
ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP
Symplectic Neural Networks for learning Generalized Hamiltonians
The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development
NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
Towards Explainable Adjudicative Variance: Quantifying Judicial Discretion via Gated Multi-Task Learning
Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation
Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning
The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans
Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
Efficient foundation decoders for fault-tolerant quantum computing
Kolmogorov Arnold networks (KAN) for aerodynamic prediction: a comparison with MLPs and GNNs
Joint Learning of Experiential Rules and Policies for Large Language Model Agents
Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage
Automating Potential-based Reward Shaping with Vision Language Model Guidance
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO
Forecasting With LLMs: Improved Generalization Through Feature Steering
Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts
LMs as Task-Specific Knowledge Bases: An Interpretability Analysis
The Geometry of Updates: Fisher Alignment at Vocabulary Scale
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
BetXplain: An Explanation-Annotated Dataset for Detecting Manipulative Betting Advertisements on Social Media
How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation
EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
Recovering Governing Equations from Solution Data: Identifiability Bounds for Linear and Nonlinear ODEs
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search
Multilingual Reasoning Cascades Need More Context
Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection
LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank
Hallucination in World Models is Predictable and Preventable
Mapping Political-Elite Networks in Europe with a Multilingual Joint Entity-Relation Extraction Pipeline
Error-Conditioned Neural Solvers
Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
google-deepmind/science-skills:GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other databases and tools.
Statistically Valid Hyperparameter Selection: From Tuning to Guarantees
Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
Probabilistic Agents in Deterministic Audits: Evaluating Multi-Agent Systems for Automated Audits Based on the German IT-Grundschutz
Taxonomy of Risks on Automated Fact-Checking Systems Considering its Propagation
MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction
Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning
Memory-Efficient Policy Libraries with Low-Rank Adaptation in Reinforcement Learning
GUI agent: Guided Exploration of User-Sensitive Screens
Cellular Predictions on the Move: What about Data?
Point Cloud Diffusion with Global and Local Reconstruction for Instance-Level 3D Anomaly Detection
Gaussian Mean Field Variational Inference can Overestimate Predictive Variance
Gradient-based inverse lithography for EUV masks via the waveguide method and a physics-informed neural operator
Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets
Bridging Spherical Black-Box Optimizers
Re-mixing Embeddings for Patient Augmentation in Data Scarce Multiple Instance Learning
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability
MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources
AI Snitches Get Glitches: Towards Evading Agentic Surveillance
Edges Before Embeddings: A Confidence-Aware Blur Gate for Vision-Language Pipelines
Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents
Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks
AutoRelAnnotator: Calibrated Model Cascades for Cost-Efficient Relevance Evaluation in Sponsored Search
Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts
Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios
Agentic System as Compressor: Quantifying System Intelligence in Bits
Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study
Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and Factorized Branch-and-Bound
Taxonomy-aware deep learning for hierarchical marine species classification in underwater imagery
Autodata: An agentic data scientist to create high quality synthetic data
Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation
Privacy Vulnerabilities of Attention Layers in Tabular Foundation Models and Protection of High-Risk Queries
Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It
TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs
Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning
AI translation of literary texts is "fine", but readers still prefer human translations
How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations
A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks
A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity
Learning Action Priors for Cross-embodiment Robot Manipulation
Measuring User's Mental Models of Speech Translation in Human-AI Collaboration
A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling
Model selection with proper scoring rules on data sets of time series
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
Task Decomposition for Efficient Annotation
Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement
Scaling Laws for Task-Specific LLM Distillation
Can Scale Save Us From Plasticity Loss in Large Language Models?
CANDLE: Character-level Arabic Noise Deduplication using Lightweight Encoder
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
Posterior Refinement: Fast Language Generation via Any-Order Flow Maps
Are We Ready For An Agent-Native Memory System?
Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce
Grad Detect: Gradient-Based Hallucination Detection in LLMs
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis
SHERLOC: Structured Diagnostic Localization for Code Repair Agents
L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System
Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
InSight: Self-Guided Skill Acquisition via Steerable VLAs
SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
Automated Residual Plot Assessment With the R Package autovi and the Shiny Application autovi.web
Social Structure Matters in 3D Human-Human Interaction Generation
SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization
MotifGen: Spatiotemporal interpolation of misaligned satellite images via multi-source generative modeling, in an application to tropical cyclones
Pigeonholing: Bad prompts hurt models to collapse and make mistakes
CALIBER: Calibrating Confidence Before and After Reasoning in Language Models
Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation
Managing Task Execution for Unknown Workloads in Batteryless IoT: A Hardware-Agnostic Evaluation
PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation
Automatic Part-of-Speech Tagging of Arabic-English Dictionary Senses through WordNet
On the Stability of Prompt Ranking in Large Language Model Evaluation
ATRIA: Adaptive Traceable ECG Reporting with Iterative Agents
Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-Imagery BCI Decoders
RE4: Transformation-aware Imitation of Object Interactions Using Manipulation Modes
Beyond Logprobs: A Multi-Signal Confidence Engine for LLM-Based Document Field Extraction
Can Aggregate Invariants Accelerate Continuous Subgraph Matching? Limits, Laws, and a Dynamic Spectral Index
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
Transformation Behavior of Images in Latent Space
MedPCFM: Improving Medical Point Cloud Completion by Integrating Point Transformers and Flow Matching
NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation
G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
UOL@IDEM at BEA 2026 Shared Task 1: Neural Fusion and Feature-Rich Modeling for L1-Aware Vocabulary Difficulty Prediction
A Fair Evaluation of Graph Foundation Models for Node Property Prediction
Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection
EERLoss: A Novel Loss Function for Training Deep Biometric Models. A Case Study in Keystroke Dynamics
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias
When CQs Go Wrong: Challenges in CQ Verification with OE-Assist
Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity
SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
Show HN: Y – A malleable coding-agent desktop app built with Electron
Abstract representational geometry supports inference in large language models
Adaptive Hard-Soft Physics-Informed Neural Networks for Robust Boundary-Constrained PDE Solving
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts
Distribution-Aware Diffusion-LLM for Robust Ultra-Long-Term Time Series Forecasting
Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems
Detecting Malicious Agent Skills in the Wild using Attention
Rethinking Object-Centric Representations for Video Dynamics Modeling
SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
Selective Time Series Forecasting via Metalearning
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
Time Series Classification through Diffeomorphic Time Warping (DiffTW)
Sublinearly Structured Deep Neural Networks Achieve Feature Learning Consistency for Compositional Functions
Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference
Self-Compacting Language Model Agents
SQLConductor: Search-to-Policy Learning for Step-wise Text-to-SQL Orchestration
Approximating velocity fields with planted attractors via Neural-ODEs for classification purposes
LangMAP: A Language-Adaptive Approach to Tokenization
Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models
Solve for the Hyperparameter, Skip the Search: Kolmogorov-Optimal Scaling Laws for Spline Regression
Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse
Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
It's Much Easier for Neural Networks to learn Game of Life Dynamics with the Right Activation Function: Polynomial Kolmogorov-Arnold Networks
A Generative Model for Closed-Loop Microsimulation of Signalized Intersections
SPIRAL: Learning to Search and Aggregate
Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking
Causal Discovery in the Era of Agents
Hedgementation = Hedgerow Segmentation: A Remote Sensing Benchmark
RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models
DiT-Reward: Generative Representations for Text-to-Image Reward Modeling
Learning Process Rewards via Success Visitation Matching for Efficient RL
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives
MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?
On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners
Can LLMs Reliably Self-Report Adversarial Prefills, and How?
Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles
PsyBridge: A Hybrid Intelligent Framework for Multi-Dimensional Mental Health Assessment and Decision Support
Semantic Browsing: Controllable Diversity for Image Generation
CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation
Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases
MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation
Any-Body Guard: Universal Safeguarding for Manipulation Policies via Action Masking
SCENIC: Semantic-Conditioned Edge-Aware Neural Framework for Structured IoT Command Generation
Encoder-Decoder Manifold Alignment for Idempotent Generation
Enhancing Protein Representation Learning via Manifold Restore Mixing
Specificity- and Calibration-Aware Breast Ultrasound Segmentation via Entropy-Guided Boundary Supervision
Benchmarking Robot Memory Under Interference
A Taxonomy of Conceptual Alignment in Human-Robot Dialogue
Reference-Free Assessment of Physical Consistency in World Model-based Video Generation
ARIA: A Causal-Aware Framework for Rescuing LLM Reasoning in Trustworthy Materials Discovery
Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers
MetaPS: Adaptive Programmatic Strategy Selection for Market Agents
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Multi-cancer detection using a computationally efficient CNN with transfer learning
Knowledge-Graph Grounding Helps LLMs Only for Out-of-Training Knowledge: A Controlled Study on Clinical Question Answering
QeHDC: Hyperdimensional Computing based on Quantum-enhanced binding and SuperClass Construction
Enhancing LLMs for Graph Tasks via Graph-aware LoRA Generation
Escaping the Variance Trap: Jacobian-Free Dynamics for Root-Finding Bilevel Optimization
Distribution-Aware Robust Bilevel Optimization: Quantile-Guided Huber Updates in Two-Timescale Stochastic Approximation
Efficient Multimodal Clinical Question Answering for Pulmonary Embolism Risk Assessment
Adaptive Recurrent Message Passing for Test Time Computing on Graphs
Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form Generation
VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows
Human and AI collaboration for pulmonary nodule segmentation
SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments
Grounded Scaling: Why Agentic AI Needs Deterministic Environments
Imagine to Ensure Safety in Hierarchical Reinforcement Learning
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do
Deep material network for homogenization of piezoelectric composites
Concept-Constrained Prompt Learning for Few-Shot CLIP Adaptation
Text2DSL: LLM-Based Code Generation for Domain-Specific Languages
Training-free Task Classification for Multi-Task Model Merging
Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline
PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
Communication Heterogeneity and Collective Consensus in Neural Cellular Automata
QBioFusion-QSAR: Morgan-Anchored Quantum Multiple Kernel Learning for Small-Data Ligand Classification
ACE-GS: Acing the Trade-off with Accurate, Compact and Efficient 3D Gaussian Splatting
Subsampling for supervised learning in reproducing kernel Hilbert spaces
Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization
Reconstructing Randomly Masked Spectra Helps DNNs Identify Discriminant Wavenumbers
Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling
NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning
Task-Differentiated Atomic Skill Expansion and Routing for Continual Learning Across Highly Heterogeneous Tasks
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
An Evaluation Framework for Text-to-Speech Voice Reconstruction
A Reward-Petri-Net Interpretation of Temporal Behavior Trees
Urban Power Grid Topology and Hierarchy Identification from Open Data
SOHET: Sequence Of Heterogeneous Events Transformer with Self-Supervised Pre-Training
Finetuning with Scientific Data Increases Hallucinations: A Multi-domain Factuality Evaluation of LLMs
NAC: Neural Action Codec for Vision-Language-Action Models
Enhancing Creativity in 3D Generative Design via a TRIZ-Inspired Text-to-CAD Framework
VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models
2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models
Dual-Attention Convolution Experts for Sparse Tensor Completion
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study
AutoRAS: Learning Robust Agentic Systems with Primitive Representations
ASCII Art Turns LLMs into VLA Controllers
Robustness Cannot be Reduced to Regularization: Studying Adversarial Training Beyond the Linear Case
Predicting High-Risk Colorectal Polyps in African Americans Using Pre-Colonoscopy Clinical Features: Machine Learning Model Development and Temporal Validation
Decoupling the Declarative from the Procedural in Vision-Language-Action Models
Breaking chains with trees: Deep learning with $\mathcal{O}(\log N)$ parallel time complexity
Towards Understanding the Power and Limits of the Muon Optimizer: A River-Valley Perspective
Backpropagating Through Simulation: Analytic Policy Gradients for Sample and Learning Efficient Differentiable Continuous Control
PeerMathDial: A Middle School Dialogue Dataset for Student Collaborative Math Problem Solving
The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection
FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving
Blogger Defeats Photographer's Copyright Claim-Sokolskyfilm vs. Messiah
You're probably using Agent Skills wrong
ningzimu/codex-ppt-skill:GPT-Image-2 PPT Generator Skill for Creating Image-Based PowerPoint Presentations in Codex and Other Skill-Compatible Agents
BuilderIO/skills:Skills for coding agents
muxuuu/serenity-skill:Serenity-inspired Agent Skill for supply-chain bottleneck stock research
0x0funky/agent-sprite-forge:Agent Skill for generating 2D sprite sheets and map, transparent PNG frames, and animated GIFs from prompts.
anysearch-ai/anysearch-skill:Unified real-time search engine skill for AI agents.
OpenBMB/PilotDeck:Task-oriented AI Agent productivity platform
earthtojake/text-to-cad:A collection of agent skills for CAD, robotics and hardware design
microsoft/SkillOpt:SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
op7418/guizang-ppt-skill:AI-agent Skill for generating polished HTML slide decks: editorial magazine and Swiss layouts, image prompts, social covers, and a WebGL/low-power presentation runtime.
Slow breathing modulates brain function and risk behavior
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
Effective Dimension Governs Generalization in Quantum Kernel Vision Models
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-trainin
Pitch Spelling Jazz Lead Sheets, Solo Transcriptions, Classical Piano and Monophonic Scores
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact
FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching
CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia
QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation
A Multi-Agent system for Multi-Objective constrained optimization
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference
Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think
The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse
Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
Statistical Properties of Training & Generalization
Recurrent neural networks approximate continuous functions
Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise
Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining
AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning
CRAX: Fast Safe Reinforcement Learning Benchmarking
Towards Modality-imbalanced Federated Graph Learning: A Data Synthesis-based Approach
Pseudo-Feature Padding: A Lightweight Defense Against False Data Injection in Power Grids
Sparsity, Superposition, and Forgetting: A Mechanistic Study of Representation Retention in Continual Learning
SSH-Net: A Deep Neural Network for Predicting Failure Time Distribution Functions under Competing Risks with Application to GPU Data
Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
Optimal Order of Multi-Agent and General Many-Body Systems
Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems
Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
Multi-Task Bayesian In-Context Learning
Toward Calibrated Mixture-of-Experts Under Distribution Shift
Optimal Deterministic Multicalibration and Omniprediction
UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning
How Transparent is DiffusionGemma?