Semantic Kernel

framework
AI 辅助整理 · 事实字段按公开来源校验

微软的 LLM 编排 SDK,面向企业与多语言,将提示、插件与规划器组合进现有应用。

内容完整度: 4/5 · 待补:related_entities

事实字段

定位
LLM 编排 SDK
主语言
C#/Python/Java
开发方
Microsoft
许可证
MIT
来源追溯: 事实字段来自公开来源整理;字段级来源元数据待继续补齐,无法稳定核验的信息不做臆测。

整理说明

本页围绕「定位、事实字段、演进时间线、关联动态」组织信息,方便快速判断 Semantic Kernel 在 AI Agent 与大模型生态中的位置。

事实型字段保持保守:未公开或无法稳定核验的信息不做臆测;关联动态保留原始来源链接,便于继续阅读。

版本演进

2026-07-08
update
BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
2026-07-08
update
Reward-Density Heuristic for Dynamic Multi-Vehicle Routing: Performance and Computational Efficiency
2026-07-08
update
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
2026-07-08
update
Scalable Perturbation Learning for Online Self-Supervised Echo State Networks
2026-07-08
update
Modeling Normal Is All You Need: Joint Latent Clustering for Anomaly Detection in Multimodal Cyber-Physical Systems
2026-07-08
update
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
2026-07-08
update
Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development
2026-07-08
update
EcoVision: AI-Powered Drone Imaging for Salt Marsh Vegetation Monitoring and Dominance Mapping
2026-07-08
update
LLM-Guided Measurement Credibility Correction for Trustworthy Industrial Process Inference
2026-07-08
update
X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models
2026-07-08
update
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
2026-07-08
update
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
2026-07-08
update
A toy framework for single and multi-agent human-AI curiosity ecosystems
2026-07-08
update
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
2026-07-08
update
Canopy: A Heterograph Foundation Model for Metabolic Engineering
2026-07-08
update
Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval
2026-07-08
update
From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition
2026-07-08
update
Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
2026-07-08
update
DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail
2026-07-08
update
Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models
2026-07-08
update
Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction
2026-07-08
update
TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting
2026-07-08
update
Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers
2026-07-08
update
Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement
2026-07-08
update
A Function-Space Dichotomy for Compositional Learning: Exponential Sub-Optimality of the Neural Tangent Kernel
2026-07-08
update
What Images Cannot Say: Language-Guided Olfactory Representation Learning
2026-07-08
update
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders
2026-07-08
update
From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b
2026-07-08
update
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
2026-07-08
update
Provable learning separation for predicting time-evolution of quantum many-body systems
2026-07-08
update
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
2026-07-08
update
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
2026-07-08
update
Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine
2026-07-08
update
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
2026-07-08
update
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment
2026-07-08
update
Graph Convolutional Attention: A Spectral Perspective on Graph Denoising and Diffusion
2026-07-07
update
Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses
2026-07-07
update
Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)
2026-07-07
update
RUFNet: Query-Guided Support Mask Refinement and Uncertainty Fusion based on Hybrid Mamba for Few-Shot Brain Tumor Segmentation
2026-07-07
update
CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion
2026-07-07
update
Toward Trustworthy Large Language Model Agents in Healthcare
2026-07-07
update
Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization
2026-07-07
update
Computing Monetary Risk Measures in Linear Time
2026-07-07
update
AIFS-SUBS: Extending Data-Driven Forecasting to Sub-Seasonal Timescales
2026-07-07
update
Choosing a parallel heterogeneous ensemble method for tabular classification
2026-07-07
update
Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters
2026-07-07
update
Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
2026-07-07
update
Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing
2026-07-07
update
PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference
2026-07-07
update
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
2026-07-07
update
ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization
2026-07-07
update
Latent Programming Horizons in Coding Agents
2026-07-07
update
Noisy-Channel Minimum Bayes Risk Decoding
2026-07-07
update
EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
2026-07-07
update
Probing Geospatial SSL Representations with Environmental Signals
2026-07-07
update
Video-based detection of cessation of breathing in pre-term infants using machine learning
2026-07-07
update
MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models
2026-07-07
update
Streaming Neural Speech Codecs through Time-Invariant Representations
2026-07-07
update
FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation
2026-07-07
update
SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment Analysis
2026-07-07
update
Target-Guided Selective Reweighting for Physics-Informed Neural Network Inverse Problems: A Transfer Learning Approach
2026-07-07
update
Untrusted Content Masking for Web Agents with Security Guarantees
2026-07-07
update
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
2026-07-07
update
Learning Only What Valid Adapters Can Express: Subspace-Constrained Adaptation Against Fine-Tuning Poisoning
2026-07-07
update
Biologically Informed Deep Neural Networks for Multi-Omic Integration, Pathway Activity Inference and Risk Stratification in Cancer
2026-07-07
update
Quantum Spectral Anomaly Detection
2026-07-07
update
Evaluating and Understanding Model Editing for Medical Vision Language Models
2026-07-07
update
How Much is Left? LLMs Linearly Encode Their Remaining Output Length
2026-07-07
update
How Far is Too Far? Defining the Distance Threshold for Verification Siamese Networks
2026-07-07
update
OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement
2026-07-07
update
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
2026-07-07
update
GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks
2026-07-07
update
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
2026-07-07
update
TabPack: Efficient Hyperparameter Ensembles for Tabular Deep Learning
2026-07-07
update
What Does a Discrete Diffusion Model Learn?
2026-07-07
update
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
2026-07-07
update
Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling
2026-07-07
update
AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes
2026-07-07
update
Environmental Drivers of Respiratory Disease: A District Level Analysis
2026-07-07
update
Transferability Between Understanding and Generation in Unified Multimodal Models
2026-07-07
update
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
2026-07-07
update
dOPSD: On-Policy Self-Distillation for Diffusion Language Models
2026-07-07
update
evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations
2026-07-07
update
Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees
2026-07-07
update
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
2026-07-07
update
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
2026-07-07
update
Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models
2026-07-07
update
Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models
2026-07-07
update
Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning
2026-07-07
update
Two Black Boxes, One Solver: Encoder Probing and Decoder Attribution for Neural Multi-Attribute VRP under Hard-Mask and Recourse Decoders
2026-07-07
update
From Interaction to Intent: Inferring User Objectives from Provenance Logs
2026-07-07
update
VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models
2026-07-07
update
Beyond travel mode: urban context shapes active mobility's mental health effects over time
2026-07-07
update
Training-Free Model Selection and Domain-Aware Score Calibration for First-Shot Anomalous Sound Detection
2026-07-07
update
Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents
2026-07-07
update
Mechanism-level routing failure in LLMs over Lean-verified algebraic structures
2026-07-07
update
CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining
2026-07-07
update
Auto: The AGI Compiler
2026-07-07
update
Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
2026-07-07
update
Explainable Novel Category Discovery in Semantic Concept Space
2026-07-07
update
MTEB-PT: A Text Embedding Benchmark for Brazilian Portuguese
2026-07-07
update
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
2026-07-07
update
LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection
2026-07-07
update
G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement
2026-07-07
update
SILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable Routing
2026-07-07
update
Machine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based Methodology
2026-07-07
update
GlaKG: A Biomarker-Centric Fundus Knowledge Graph for Explainable Glaucoma Diagnosis and Risk Assessment
2026-07-07
update
URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment
2026-07-07
update
PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection
2026-07-06
update
yetone/native-feel-skill:An Agent Skill for designing cross-platform desktop apps that feel native — distilled from Raycast's 2.0 deep-dive and reverse engineering of Raycast Beta.app. Eight architectural tenets, four-layer architecture, WebKit/WebView2 survival guide, 75-item ship audit.
2026-07-04
update
rpamis/comet:Comet: agent skill harness phase-guarded automation from idea to archive
2026-07-04
update
Show HN: TaskPeace – a task queue my AI coding agents pull work from over MCP
2026-07-03
update
Coding-agents can replicate scientific machine learning papers
2026-07-03
update
Probing Chemical Language Models: Effects of Pre-training and Fine-tuning
2026-07-03
update
Dynamic Neural Graph Encoding of Inference Processes in Deep Weight Space
2026-07-03
update
Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation
2026-07-03
update
RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation
2026-07-03
update
UA-ChatDev: Uncertainty-Aware Multi-Agent Collaboration for Reliable Software Development
2026-07-03
update
An Optimisation Framework for the Well-Conditioned Training of Physics-Informed Neural Networks
2026-07-03
update
Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond
2026-07-03
update
The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
2026-07-03
update
Prediction Sets for Counterfactual Decisions: Coverage, Optimality, and Conformal Prediction
2026-07-03
update
Criticality-Based Guard Rail Validation for AI Agent Decisions in Autonomous Telecom Networks
2026-07-03
update
An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility
2026-07-03
update
Aggregation with Exponential Weights is Optimal in Expectation
2026-07-03
update
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
2026-07-03
update
HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures
2026-07-03
update
Dendritic In-Context Learning in a Single-Layer Spiking Neural Network
2026-07-03
update
Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics
2026-07-03
update
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
2026-07-03
update
SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces
2026-07-03
update
GAP-GDRNet: Geometry-Aware Monocular Visual Pose Sensing on a Single-Target Synthetic Spacecraft Dataset
2026-07-03
update
Stable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory Gates
2026-07-03
update
VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval
2026-07-03
update
Know Your Source: A Public Knowledge Store for Media Background Checks
2026-07-03
update
Bringing Agentic Search to Earth Observation Data Discovery
2026-07-03
update
Steerability via constraints: a substrate for scalable oversight of coding agents
2026-07-03
update
ACID: Action Consistency via Inverse Dynamics for Planning with World Models
2026-07-03
update
Object-centric LeJEPA
2026-07-03
update
Q-GAIN: A Python Package for Machine Learning and Physically Informed Analysis Applications
2026-07-03
update
The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing
2026-07-03
update
Neuron-Aware Active Few-Shot Learning for LLMs
2026-07-03
update
WorldSample: Closed-loop Real-robot RL with World Modelling
2026-07-03
update
Extreme Adaptive Transformer for Time Series Forecasting
2026-07-03
update
Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID Data
2026-07-03
update
Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation
2026-07-03
update
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
2026-07-03
update
Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning
2026-07-03
update
Controllable Sim Agents with Behavior Latents
2026-07-03
update
Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas
2026-07-03
update
What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates
2026-07-03
update
Online Safety Monitoring for LLMs
2026-07-03
update
LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
2026-07-03
update
Show HN: QUALITY.md – open format/specification, agent skill, and CLI
2026-07-02
update
Show HN: Mail Memories – A desktop app to rescue photos from Gmail
2026-07-02
update
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
2026-07-02
update
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models
2026-07-02
update
Dynamic Bidirectional Pattern Memory: A Production-Scale Empirical Characterisation of Inference-Time Gating in Clinical NLP
2026-07-02
update
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
2026-07-02
update
Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm
2026-07-02
update
Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions
2026-07-02
update
Aionoscope: Debugging Latent-State Accessibility in Time-Series Representations
2026-07-02
update
Bridging Quantum Computing Paradigms toward Semiconductor Yield: A Controlled CV-versus-DV Comparison on Wafer-Map Defect Classification
2026-07-02
update
TRCGL-Net: A Long-Tailed Multi-Label Chest X-Ray Classification Framework with Generative Data Augmentation and Label Co-Occurrence Modeling
2026-07-02
update
Foundation Models vs. Radiomics for Lung Computed Tomography: A Benchmark of Feature Extractors, Classification Heads, and Segmentation Choices
2026-07-02
update
Understanding Large Language Models
2026-07-02
update
Evidence-Supported Credit Risk Report Generation Using News-Centric Financial Knowledge Graphs
2026-07-02
update
Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework
2026-07-02
update
EchoRisk: A Multicentre Echocardiography Dataset and Benchmark for Cardio-Oncology
2026-07-02
update
Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
2026-07-02
update
GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache
2026-07-02
update
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
2026-07-02
update
Message Passing Enables Efficient Reasoning
2026-07-02
update
Balancing Expressivity and Learnability in Quantum Kernel Bandit Optimization
2026-07-02
update
Group-invariant Coresets for Data-efficient Active Learning
2026-07-02
update
SynLaD: Latent Diffusion for Generating Synthesizable Molecules Conditioned on 3D Pharmacophore Profiles
2026-07-02
update
FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement
2026-07-02
update
GAIA: Geometry-Adaptive Operator Learning for Forward and Inverse Problems
2026-07-02
update
Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search
2026-07-02
update
A Lightweight Self-Supervised Learning Framework for Multivariate Time Series using Hierarchical-JEPA on ECG Data
2026-07-02
update
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
2026-07-02
update
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
2026-07-02
update
Decision-Aware Training for Sample-Based Generative Models
2026-07-02
update
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations
2026-07-02
update
Optimal Resource Utilization for Autonomous Laboratory Orchestrators
2026-07-02
update
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
2026-07-02
update
FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model
2026-07-02
update
The State-Prediction Separation Hypothesis
2026-07-02
update
Language-Critique Imitation Learning from Suboptimal Demonstrations
2026-07-02
update
Measuring the Gap Between Human and LLM Research Ideas
2026-07-02
update
Show HN: AnalystAIPack – 118 runnable agent skills for malware analysis and RE
2026-07-01
update
cloudflare/security-audit-skill:A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
2026-07-01
update
When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection
2026-07-01
update
Overview of the TalentCLEF 2026: Skill and Job Title Intelligence for Human Capital Management
2026-07-01
update
ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping
2026-07-01
update
Diffusing Blame: Task-Dependent Credit Assignment in Biologically Plausible Dual-Stream Networks
2026-07-01
update
Nonlinearity-Aware LoRA: Structured Gate Adaptation under Low-Rank Constraints
2026-07-01
update
Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian
2026-07-01
update
STEB: Style Text Embedding Benchmark
2026-07-01
update
JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering
2026-07-01
update
A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols
2026-07-01
update
Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
2026-07-01
update
SpikeLogBERT: Energy-Efficient Log Parsing Using Spiking Transformer Networks
2026-07-01
update
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
2026-07-01
update
Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision
2026-07-01
update
Relational and Sequential Conformal Inference for Energy Time Series over Graphs via Foundation Models
2026-07-01
update
Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models
2026-07-01
update
Low-dimensional topology of deep neural networks
2026-07-01
update
Accelerating Conformal Prediction via Approximate Leave-One-Out
2026-07-01
update
Semantic Leakage and Privacy Preservation in Relay-Assisted Semantic Communications
2026-07-01
update
TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models
2026-07-01
update
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
2026-07-01
update
Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization
2026-07-01
update
Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA
2026-07-01
update
CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation
2026-07-01
update
Scalable Behaviour Cloning on Browser Using via Skill Distillation
2026-07-01
update
FedLAB: Traceable Semantic Codebooks for Federated Multimodal Graph Foundation Learning
2026-07-01
update
Automated Background Swapping for Robustness against Spurious Backgrounds
2026-07-01
update
SemRF: A Semantic Reference Frame for Residual-Stream Dynamics in Language Models
2026-07-01
update
Generative Skill Composition for LLM Agents
2026-07-01
update
AdaJEPA: An Adaptive Latent World Model
2026-07-01
update
Freeform Preference Learning for Robotic Manipulation
2026-07-01
update
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors
2026-07-01
update
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
2026-07-01
update
Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision
2026-07-01
update
Spatial Reasoning via Modality Switching Between Language and Symbolic Representation
2026-07-01
update
From Idea to Prototype in an Afternoon: Scaffolded, AI-Assisted Rapid VA Prototyping
2026-07-01
update
Safe Online Learning via Smooth Safety-Structured Policy Composition
2026-07-01
update
PGUDA: Pressure-Guided Unsupervised Domain Adaptation with Cross-Modal Knowledge Distillation for sEMG-Based Gesture Recognition
2026-07-01
update
Dualformer: Efficient Feature Extractor for Complex-valued Blind Communication Signal Analysis
2026-07-01
update
Stage-Transition Dense Reward Modeling for Reinforcement Learning
2026-07-01
update
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
2026-07-01
update
World-Model Collapse as a Phase Transition
2026-07-01
update
Xiaomi-GUI-0 Technical Report
2026-07-01
update
DA-Studio: An Agentic System for End-to-End Data Analysis
2026-07-01
update
Ask the World Before Acting: Budgeted Environment Probing for World-Model Calibration
2026-07-01
update
CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes
2026-07-01
update
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
2026-07-01
update
Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models
2026-07-01
update
CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market
2026-07-01
update
Team MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline Dynamics
2026-07-01
update
Von Mises Based Uncertainty Quantification for Closely Spaced Automotive Radar Targets
2026-07-01
update
Surprise as a Signal for Plasticity and Metacognition
2026-07-01
update
Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models
2026-07-01
update
Design and Implementation of Agentic Orchestrations and Orchestration of Agents
2026-07-01
update
RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference
2026-07-01
update
On the Convergence of Self-Improving Online LLM Alignment
2026-07-01
update
Beyond the Expressivity-Trainability Paradox: A Dynamical Lie Algebra Perspective on Navigating Barren Plateaus in Quantum Machine Learning
2026-07-01
update
ACE: Pluggable Adaptive Context Elasticizer across Agents
2026-07-01
update
Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning
2026-07-01
update
DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers
2026-07-01
update
Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning
2026-07-01
update
Preserve the Hard, Regenerate the Rest: Uncertainty-Guided Synthetic Training Data Augmentation with Diffusion Models
2026-07-01
update
A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems
2026-07-01
update
FARS: A Fully Automated Research System Deployed at Scale
2026-06-30
update
Extrapolating from Regularised Solutions for Solving Ill-Conditioned Linear Systems in Machine Learning
2026-06-30
update
BayesEvolve: Explicit Belief States for Autonomous Scientific Discovery
2026-06-30
update
FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks
2026-06-30
update
REAR: Test-time Preference Realignment through Reward Decomposition
2026-06-30
update
Early Cue Precision Shapes Visual Shortcut Learning in Controlled Cue-Manipulation Benchmarks
2026-06-30
update
Residual-Guided Expert Specialization for Incomplete Multimodal Learning
2026-06-30
update
OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL
2026-06-30
update
On the Vulnerability of Parameter-Level Defenses to Model Merging
2026-06-30
update
Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data
2026-06-30
update
Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation
2026-06-30
update
Scalar Representations of Neural Network Training Dynamics
2026-06-30
update
Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding
2026-06-30
update
Beyond IID: How General Are Tabular Foundation Models, Really?
2026-06-30
update
Can LLMs Rank? A Tale of Triads and Triage
2026-06-30
update
Beyond Point Estimates for Glaucoma Visual Field Forecasting with Diffusion Models
2026-06-30
update
Translating Natural Language to Strategic Temporal Specifications via LLMs
2026-06-30
update
The FIL Hypothesis: Inductive Biases Help with Kernel Engineering
2026-06-30
update
On the Faithfulness of Post-Hoc Concept Bottleneck Models
2026-06-30
update
Discovering Collaboration from Novelty: Random Network Distillation for Clustered Federated Learning
2026-06-30
update
$μ$Flow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors
2026-06-30
update
Entity Binding Failures in Tool-Augmented Agents
2026-06-30
update
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
2026-06-30
update
Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?
2026-06-30
update
Convergence of Continual Learning in Homogeneous Deep Networks
2026-06-30
update
The Human Creativity Benchmark
2026-06-30
update
A Multi-task Mixture of Experts Framework for Malware Classification, Packing Detection, and Family Attribution
2026-06-30
update
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
2026-06-30
update
The Fundamental Limits of Valid Transport Map Estimation
2026-06-30
update
Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization
2026-06-30
update
Uncertainty-Aware Generation and Decision-Making Under Ambiguity
2026-06-30
update
Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection
2026-06-30
update
Wireless Backdoor Attack and Defense for Semantic Communications over Multiple Access Channel
2026-06-30
update
Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms
2026-06-30
update
DOPD: Dual On-policy Distillation
2026-06-30
update
GROW$^2$: Grounding Which and Where for Robot Tool Use
2026-06-30
update
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
2026-06-30
update
Exploiting Local Flatness for Efficient Out-of-Distribution Detection
2026-06-30
update
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
2026-06-30
update
SWE-Together: Evaluating Coding Agents in Interactive User Sessions
2026-06-30
update
DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation
2026-06-30
update
RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines
2026-06-30
update
NeuReasoner: Theory-grounded Mapping of Reasoning Elicitation Boundaries
2026-06-30
update
First-Order Temporal Logic Tensor Networks
2026-06-30
update
Exploration and Online Transfer with Behavioral Foundation Models
2026-06-30
update
Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping
2026-06-30
update
AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills
2026-06-30
update
T3R: Deeper Test-Time Adaptation for Graph Neural Networks via Gradient Rotation
2026-06-30
update
Consensus Clustering of Free-Viewing Gaze Data: New Insights into Human-Information Interaction
2026-06-30
update
Heads, Not Backbones: Output Heads Dominate Architectures on Fat-Tailed Returns
2026-06-30
update
Building Multi-Task Agentic LLMs via Two-Phase Distillation
2026-06-30
update
Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance
2026-06-30
update
Data-Driven Energy-Based Learning via Gibbs Measures on Hierarchical Structures
2026-06-30
update
Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management
2026-06-30
update
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
2026-06-30
update
Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts
2026-06-30
update
Efficient Retrieval-Augmented Generation via Token Co-occurrence Graphs
2026-06-30
update
Temporal Feature Extractors in EEG Foundation Models: A Controlled Comparison Including a Pretrained Time-Series Model
2026-06-30
update
Automating the Design of Embodied AgentArchitectures
2026-06-30
update
SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance
2026-06-30
update
Gravitational Duals from Equations of State II: Large Hierarchies and False Vacua
2026-06-30
update
Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters
2026-06-30
update
Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching
2026-06-30
update
DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks
2026-06-30
update
Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark
2026-06-30
update
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
2026-06-30
update
DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning
2026-06-30
update
From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent
2026-06-30
update
Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data
2026-06-30
update
Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion
2026-06-30
update
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning
2026-06-30
update
B3O: Scalable Boltzmann Batch Bayesian Optimization
2026-06-30
update
A Distributionally Robust Framework for Learned Reconstructions in Inverse Problems
2026-06-30
update
CaresAI at CT-DEB26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Models
2026-06-30
update
Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration
2026-06-30
update
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
2026-06-30
update
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
2026-06-30
update
KnowsTFM: Knowledge-Informed Fine-Tuning of Small Tabular Foundation Models
2026-06-30
update
Defending Against Harmful Supervision Hidden in Benign Samples
2026-06-30
update
Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation
2026-06-30
update
PromptGNN-sim: Deep Fusion and Alignment of GNN and LLMs for Text-Attributed Graph Learning
2026-06-30
update
ManimAgent: Self-Evolving Multimodal Agents for Visual Education
2026-06-30
update
Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents
2026-06-30
update
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
2026-06-30
update
Forward-Future/loopy:A library of practical AI-agent loops and an installable skill for finding, adapting, and designing repeatable agent workflows.
2026-06-29
update
Applicability of memorization indicators for early spotting of overfitting while recalibrating sEMG-decoders on low sample sizes
2026-06-29
update
S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
2026-06-29
update
SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
2026-06-29
update
A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts
2026-06-29
update
Mosaic: A Benchmark Suite for Differentiable Physics Solvers
2026-06-29
update
Home3D 1.0: A High-Fidelity Image-to-3D Asset Generation System for Interior Design
2026-06-29
update
Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy
2026-06-29
update
AI Persuasive Framing in Collective Dilemmas
2026-06-29
update
An Empirical Analysis of Factual Errors in Human-Written Text and its Application
2026-06-29
update
Reasoning Beyond Prediction: From Data-Driven to Causal Software Engineering
2026-06-29
update
ToxiREX: A Dataset on Toxic REasoning in ConteXt
2026-06-29
update
Benchmarking on Tasks That Matter: Dataset Selection for Preserving Model Rankings
2026-06-29
update
Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection
2026-06-29
update
Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA
2026-06-29
update
ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
2026-06-29
update
Single and Multi Truth Data Fusion using Large Language Models
2026-06-29
update
JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
2026-06-29
update
OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
2026-06-29
update
Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training
2026-06-29
update
When One Adapter Speaks for Many: Discovering Low-Rank Redundancy in Continual Fine-Tuning
2026-06-29
update
Beyond Sparse Supervision: Diffusion-Guided Learning for Few-Shot Graph Fraud Detection
2026-06-29
update
Toward Robust In-Context Segmentation via Concept Guidance
2026-06-29
update
Regularized Reward-Punishment Reinforcement Learning
2026-06-29
update
EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
2026-06-29
update
CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association
2026-06-29
update
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
2026-06-29
update
The Remittance Blueprint: Data-driven Intelligence for Sri Lanka
2026-06-29
update
COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives
2026-06-29
update
Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software
2026-06-29
update
Bridging Ab Initio Symmetries and Global Nuclear Masses with Interpretable Neural Networks
2026-06-29
update
Democratic ICAI: Debating Our Way to Steering Principles from Preferences
2026-06-29
update
VGB for Masked Diffusion Model: Efficient Test-time Scaling for Reward Satisfaction and Sample Editing
2026-06-29
update
Which Nash Equilibrium? Solver-Dependent Selection on Zero-Sum Nash Polytopes
2026-06-29
update
DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand
2026-06-29
update
AI boom risks global financial crash, warn central bankers
2026-06-27
update
Task Failed Successfully: Saturating NIC and Disk Bandwidth
2026-06-27
update
Aircraft crashes into Beijing's tallest skyscraper, triggering evacuations
2026-06-26
update
Einstein World Models
2026-06-26
update
XMSE-Aware Adaptive Empirical Bayes Estimation
2026-06-26
update
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics
2026-06-26
update
Decision-Aligned Evaluation of Uncertainty Quantification
2026-06-26
update
Event-Aware Instructed Assistant for Referring Video Segmentation
2026-06-26
update
On-board Remote-Sensing Foundation Models for Unsupervised Change Detection of Disaster Events
2026-06-26
update
ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP
2026-06-26
update
Symplectic Neural Networks for learning Generalized Hamiltonians
2026-06-26
update
The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development
2026-06-26
update
NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
2026-06-26
update
Towards Explainable Adjudicative Variance: Quantifying Judicial Discretion via Gated Multi-Task Learning
2026-06-26
update
Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation
2026-06-26
update
Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning
2026-06-26
update
The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans
2026-06-26
update
Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions
2026-06-26
update
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
2026-06-26
update
Efficient foundation decoders for fault-tolerant quantum computing
2026-06-26
update
Kolmogorov Arnold networks (KAN) for aerodynamic prediction: a comparison with MLPs and GNNs
2026-06-26
update
Joint Learning of Experiential Rules and Policies for Large Language Model Agents
2026-06-26
update
Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
2026-06-26
update
OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
2026-06-26
update
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
2026-06-26
update
Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
2026-06-26
update
RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage
2026-06-26
update
Automating Potential-based Reward Shaping with Vision Language Model Guidance
2026-06-26
update
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
2026-06-26
update
A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO
2026-06-26
update
Forecasting With LLMs: Improved Generalization Through Feature Steering
2026-06-26
update
Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
2026-06-26
update
Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts
2026-06-26
update
LMs as Task-Specific Knowledge Bases: An Interpretability Analysis
2026-06-26
update
The Geometry of Updates: Fisher Alignment at Vocabulary Scale
2026-06-26
update
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
2026-06-26
update
BetXplain: An Explanation-Annotated Dataset for Detecting Manipulative Betting Advertisements on Social Media
2026-06-26
update
How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation
2026-06-26
update
EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
2026-06-26
update
Recovering Governing Equations from Solution Data: Identifiability Bounds for Linear and Nonlinear ODEs
2026-06-26
update
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
2026-06-26
update
Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search
2026-06-26
update
Multilingual Reasoning Cascades Need More Context
2026-06-26
update
Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection
2026-06-26
update
LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank
2026-06-26
update
Hallucination in World Models is Predictable and Preventable
2026-06-26
update
Mapping Political-Elite Networks in Europe with a Multilingual Joint Entity-Relation Extraction Pipeline
2026-06-26
update
Error-Conditioned Neural Solvers
2026-06-26
update
Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
2026-06-26
update
google-deepmind/science-skills:GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other databases and tools.
2026-06-25
update
Statistically Valid Hyperparameter Selection: From Tuning to Guarantees
2026-06-25
update
Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
2026-06-25
update
Probabilistic Agents in Deterministic Audits: Evaluating Multi-Agent Systems for Automated Audits Based on the German IT-Grundschutz
2026-06-25
update
Taxonomy of Risks on Automated Fact-Checking Systems Considering its Propagation
2026-06-25
update
MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction
2026-06-25
update
Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning
2026-06-25
update
Memory-Efficient Policy Libraries with Low-Rank Adaptation in Reinforcement Learning
2026-06-25
update
GUI agent: Guided Exploration of User-Sensitive Screens
2026-06-25
update
Cellular Predictions on the Move: What about Data?
2026-06-25
update
Point Cloud Diffusion with Global and Local Reconstruction for Instance-Level 3D Anomaly Detection
2026-06-25
update
Gaussian Mean Field Variational Inference can Overestimate Predictive Variance
2026-06-25
update
Gradient-based inverse lithography for EUV masks via the waveguide method and a physics-informed neural operator
2026-06-25
update
Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets
2026-06-25
update
Bridging Spherical Black-Box Optimizers
2026-06-25
update
Re-mixing Embeddings for Patient Augmentation in Data Scarce Multiple Instance Learning
2026-06-25
update
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
2026-06-25
update
Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability
2026-06-25
update
MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources
2026-06-25
update
AI Snitches Get Glitches: Towards Evading Agentic Surveillance
2026-06-25
update
Edges Before Embeddings: A Confidence-Aware Blur Gate for Vision-Language Pipelines
2026-06-25
update
Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents
2026-06-25
update
Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks
2026-06-25
update
AutoRelAnnotator: Calibrated Model Cascades for Cost-Efficient Relevance Evaluation in Sponsored Search
2026-06-25
update
Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts
2026-06-25
update
Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
2026-06-25
update
SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios
2026-06-25
update
Agentic System as Compressor: Quantifying System Intelligence in Bits
2026-06-25
update
Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study
2026-06-25
update
Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and Factorized Branch-and-Bound
2026-06-25
update
Weave of Formal Thought
2026-06-25
update
Taxonomy-aware deep learning for hierarchical marine species classification in underwater imagery
2026-06-25
update
Autodata: An agentic data scientist to create high quality synthetic data
2026-06-25
update
Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
2026-06-25
update
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation
2026-06-25
update
Privacy Vulnerabilities of Attention Layers in Tabular Foundation Models and Protection of High-Risk Queries
2026-06-25
update
Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It
2026-06-25
update
TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs
2026-06-25
update
Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning
2026-06-25
update
AI translation of literary texts is "fine", but readers still prefer human translations
2026-06-25
update
How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations
2026-06-25
update
A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks
2026-06-25
update
A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding
2026-06-25
update
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
2026-06-25
update
On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity
2026-06-25
update
Learning Action Priors for Cross-embodiment Robot Manipulation
2026-06-24
update
Measuring User's Mental Models of Speech Translation in Human-AI Collaboration
2026-06-24
update
A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling
2026-06-24
update
Model selection with proper scoring rules on data sets of time series
2026-06-24
update
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
2026-06-24
update
Task Decomposition for Efficient Annotation
2026-06-24
update
Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement
2026-06-24
update
Scaling Laws for Task-Specific LLM Distillation
2026-06-24
update
Can Scale Save Us From Plasticity Loss in Large Language Models?
2026-06-24
update
CANDLE: Character-level Arabic Noise Deduplication using Lightweight Encoder
2026-06-24
update
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
2026-06-24
update
Posterior Refinement: Fast Language Generation via Any-Order Flow Maps
2026-06-24
update
Are We Ready For An Agent-Native Memory System?
2026-06-24
update
Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce
2026-06-24
update
Grad Detect: Gradient-Based Hallucination Detection in LLMs
2026-06-24
update
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
2026-06-24
update
OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis
2026-06-24
update
SHERLOC: Structured Diagnostic Localization for Code Repair Agents
2026-06-24
update
L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models
2026-06-24
update
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System
2026-06-24
update
Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models
2026-06-24
update
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
2026-06-24
update
InSight: Self-Guided Skill Acquisition via Steerable VLAs
2026-06-24
update
SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
2026-06-24
update
Automated Residual Plot Assessment With the R Package autovi and the Shiny Application autovi.web
2026-06-24
update
Social Structure Matters in 3D Human-Human Interaction Generation
2026-06-24
update
SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization
2026-06-24
update
MotifGen: Spatiotemporal interpolation of misaligned satellite images via multi-source generative modeling, in an application to tropical cyclones
2026-06-24
update
Pigeonholing: Bad prompts hurt models to collapse and make mistakes
2026-06-24
update
CALIBER: Calibrating Confidence Before and After Reasoning in Language Models
2026-06-24
update
Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation
2026-06-24
update
Managing Task Execution for Unknown Workloads in Batteryless IoT: A Hardware-Agnostic Evaluation
2026-06-24
update
PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation
2026-06-24
update
Automatic Part-of-Speech Tagging of Arabic-English Dictionary Senses through WordNet
2026-06-24
update
On the Stability of Prompt Ranking in Large Language Model Evaluation
2026-06-24
update
ATRIA: Adaptive Traceable ECG Reporting with Iterative Agents
2026-06-24
update
Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-Imagery BCI Decoders
2026-06-24
update
RE4: Transformation-aware Imitation of Object Interactions Using Manipulation Modes
2026-06-24
update
Beyond Logprobs: A Multi-Signal Confidence Engine for LLM-Based Document Field Extraction
2026-06-24
update
Can Aggregate Invariants Accelerate Continuous Subgraph Matching? Limits, Laws, and a Dynamic Spectral Index
2026-06-24
update
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
2026-06-24
update
Transformation Behavior of Images in Latent Space
2026-06-24
update
MedPCFM: Improving Medical Point Cloud Completion by Integrating Point Transformers and Flow Matching
2026-06-24
update
NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation
2026-06-24
update
G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
2026-06-24
update
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
2026-06-24
update
UOL@IDEM at BEA 2026 Shared Task 1: Neural Fusion and Feature-Rich Modeling for L1-Aware Vocabulary Difficulty Prediction
2026-06-24
update
A Fair Evaluation of Graph Foundation Models for Node Property Prediction
2026-06-24
update
Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation
2026-06-24
update
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
2026-06-24
update
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
2026-06-24
update
Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection
2026-06-24
update
EERLoss: A Novel Loss Function for Training Deep Biometric Models. A Case Study in Keystroke Dynamics
2026-06-24
update
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
2026-06-24
update
To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias
2026-06-24
update
When CQs Go Wrong: Challenges in CQ Verification with OE-Assist
2026-06-24
update
Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity
2026-06-24
update
SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation
2026-06-24
update
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
2026-06-24
update
Show HN: Y – A malleable coding-agent desktop app built with Electron
2026-06-23
update
Abstract representational geometry supports inference in large language models
2026-06-23
update
Adaptive Hard-Soft Physics-Informed Neural Networks for Robust Boundary-Constrained PDE Solving
2026-06-23
update
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts
2026-06-23
update
Distribution-Aware Diffusion-LLM for Robust Ultra-Long-Term Time Series Forecasting
2026-06-23
update
Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems
2026-06-23
update
Detecting Malicious Agent Skills in the Wild using Attention
2026-06-23
update
Rethinking Object-Centric Representations for Video Dynamics Modeling
2026-06-23
update
SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
2026-06-23
update
Selective Time Series Forecasting via Metalearning
2026-06-23
update
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
2026-06-23
update
Time Series Classification through Diffeomorphic Time Warping (DiffTW)
2026-06-23
update
Sublinearly Structured Deep Neural Networks Achieve Feature Learning Consistency for Compositional Functions
2026-06-23
update
Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference
2026-06-23
update
Self-Compacting Language Model Agents
2026-06-23
update
SQLConductor: Search-to-Policy Learning for Step-wise Text-to-SQL Orchestration
2026-06-23
update
Approximating velocity fields with planted attractors via Neural-ODEs for classification purposes
2026-06-23
update
LangMAP: A Language-Adaptive Approach to Tokenization
2026-06-23
update
Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models
2026-06-23
update
Solve for the Hyperparameter, Skip the Search: Kolmogorov-Optimal Scaling Laws for Spline Regression
2026-06-23
update
Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse
2026-06-23
update
Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
2026-06-23
update
It's Much Easier for Neural Networks to learn Game of Life Dynamics with the Right Activation Function: Polynomial Kolmogorov-Arnold Networks
2026-06-23
update
A Generative Model for Closed-Loop Microsimulation of Signalized Intersections
2026-06-23
update
SPIRAL: Learning to Search and Aggregate
2026-06-23
update
Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking
2026-06-23
update
Causal Discovery in the Era of Agents
2026-06-23
update
Hedgementation = Hedgerow Segmentation: A Remote Sensing Benchmark
2026-06-23
update
RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models
2026-06-23
update
DiT-Reward: Generative Representations for Text-to-Image Reward Modeling
2026-06-23
update
Learning Process Rewards via Success Visitation Matching for Efficient RL
2026-06-23
update
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
2026-06-23
update
Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives
2026-06-23
update
MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?
2026-06-23
update
On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners
2026-06-23
update
Can LLMs Reliably Self-Report Adversarial Prefills, and How?
2026-06-23
update
Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles
2026-06-23
update
PsyBridge: A Hybrid Intelligent Framework for Multi-Dimensional Mental Health Assessment and Decision Support
2026-06-23
update
Semantic Browsing: Controllable Diversity for Image Generation
2026-06-23
update
CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation
2026-06-23
update
Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases
2026-06-23
update
MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation
2026-06-23
update
Any-Body Guard: Universal Safeguarding for Manipulation Policies via Action Masking
2026-06-23
update
SCENIC: Semantic-Conditioned Edge-Aware Neural Framework for Structured IoT Command Generation
2026-06-23
update
Encoder-Decoder Manifold Alignment for Idempotent Generation
2026-06-23
update
Enhancing Protein Representation Learning via Manifold Restore Mixing
2026-06-23
update
Specificity- and Calibration-Aware Breast Ultrasound Segmentation via Entropy-Guided Boundary Supervision
2026-06-23
update
Benchmarking Robot Memory Under Interference
2026-06-23
update
A Taxonomy of Conceptual Alignment in Human-Robot Dialogue
2026-06-23
update
Reference-Free Assessment of Physical Consistency in World Model-based Video Generation
2026-06-23
update
ARIA: A Causal-Aware Framework for Rescuing LLM Reasoning in Trustworthy Materials Discovery
2026-06-23
update
Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers
2026-06-23
update
MetaPS: Adaptive Programmatic Strategy Selection for Market Agents
2026-06-23
update
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
2026-06-23
update
Multi-cancer detection using a computationally efficient CNN with transfer learning
2026-06-23
update
Knowledge-Graph Grounding Helps LLMs Only for Out-of-Training Knowledge: A Controlled Study on Clinical Question Answering
2026-06-23
update
QeHDC: Hyperdimensional Computing based on Quantum-enhanced binding and SuperClass Construction
2026-06-23
update
Enhancing LLMs for Graph Tasks via Graph-aware LoRA Generation
2026-06-23
update
Escaping the Variance Trap: Jacobian-Free Dynamics for Root-Finding Bilevel Optimization
2026-06-23
update
Distribution-Aware Robust Bilevel Optimization: Quantile-Guided Huber Updates in Two-Timescale Stochastic Approximation
2026-06-23
update
Efficient Multimodal Clinical Question Answering for Pulmonary Embolism Risk Assessment
2026-06-23
update
Adaptive Recurrent Message Passing for Test Time Computing on Graphs
2026-06-23
update
Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form Generation
2026-06-23
update
VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows
2026-06-23
update
Human and AI collaboration for pulmonary nodule segmentation
2026-06-23
update
SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments
2026-06-23
update
Grounded Scaling: Why Agentic AI Needs Deterministic Environments
2026-06-23
update
Imagine to Ensure Safety in Hierarchical Reinforcement Learning
2026-06-23
update
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do
2026-06-23
update
Deep material network for homogenization of piezoelectric composites
2026-06-23
update
Concept-Constrained Prompt Learning for Few-Shot CLIP Adaptation
2026-06-23
update
Text2DSL: LLM-Based Code Generation for Domain-Specific Languages
2026-06-23
update
Training-free Task Classification for Multi-Task Model Merging
2026-06-23
update
Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline
2026-06-23
update
PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
2026-06-23
update
Communication Heterogeneity and Collective Consensus in Neural Cellular Automata
2026-06-23
update
QBioFusion-QSAR: Morgan-Anchored Quantum Multiple Kernel Learning for Small-Data Ligand Classification
2026-06-23
update
ACE-GS: Acing the Trade-off with Accurate, Compact and Efficient 3D Gaussian Splatting
2026-06-23
update
Subsampling for supervised learning in reproducing kernel Hilbert spaces
2026-06-23
update
Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization
2026-06-23
update
Reconstructing Randomly Masked Spectra Helps DNNs Identify Discriminant Wavenumbers
2026-06-23
update
Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling
2026-06-23
update
NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning
2026-06-23
update
Task-Differentiated Atomic Skill Expansion and Routing for Continual Learning Across Highly Heterogeneous Tasks
2026-06-23
update
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
2026-06-23
update
An Evaluation Framework for Text-to-Speech Voice Reconstruction
2026-06-23
update
A Reward-Petri-Net Interpretation of Temporal Behavior Trees
2026-06-23
update
Urban Power Grid Topology and Hierarchy Identification from Open Data
2026-06-23
update
SOHET: Sequence Of Heterogeneous Events Transformer with Self-Supervised Pre-Training
2026-06-23
update
Finetuning with Scientific Data Increases Hallucinations: A Multi-domain Factuality Evaluation of LLMs
2026-06-23
update
NAC: Neural Action Codec for Vision-Language-Action Models
2026-06-23
update
Enhancing Creativity in 3D Generative Design via a TRIZ-Inspired Text-to-CAD Framework
2026-06-23
update
VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models
2026-06-23
update
2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models
2026-06-23
update
Dual-Attention Convolution Experts for Sparse Tensor Completion
2026-06-23
update
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study
2026-06-23
update
AutoRAS: Learning Robust Agentic Systems with Primitive Representations
2026-06-23
update
ASCII Art Turns LLMs into VLA Controllers
2026-06-23
update
Robustness Cannot be Reduced to Regularization: Studying Adversarial Training Beyond the Linear Case
2026-06-23
update
Predicting High-Risk Colorectal Polyps in African Americans Using Pre-Colonoscopy Clinical Features: Machine Learning Model Development and Temporal Validation
2026-06-23
update
Decoupling the Declarative from the Procedural in Vision-Language-Action Models
2026-06-23
update
Breaking chains with trees: Deep learning with $\mathcal{O}(\log N)$ parallel time complexity
2026-06-23
update
Towards Understanding the Power and Limits of the Muon Optimizer: A River-Valley Perspective
2026-06-23
update
Backpropagating Through Simulation: Analytic Policy Gradients for Sample and Learning Efficient Differentiable Continuous Control
2026-06-23
update
PeerMathDial: A Middle School Dialogue Dataset for Student Collaborative Math Problem Solving
2026-06-23
update
The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection
2026-06-23
update
FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving
2026-06-23
update
Blogger Defeats Photographer's Copyright Claim-Sokolskyfilm vs. Messiah
2026-06-22
update
You're probably using Agent Skills wrong
2026-06-22
update
ningzimu/codex-ppt-skill:GPT-Image-2 PPT Generator Skill for Creating Image-Based PowerPoint Presentations in Codex and Other Skill-Compatible Agents
2026-06-21
update
BuilderIO/skills:Skills for coding agents
2026-06-21
update
muxuuu/serenity-skill:Serenity-inspired Agent Skill for supply-chain bottleneck stock research
2026-06-21
update
0x0funky/agent-sprite-forge:Agent Skill for generating 2D sprite sheets and map, transparent PNG frames, and animated GIFs from prompts.
2026-06-21
update
anysearch-ai/anysearch-skill:Unified real-time search engine skill for AI agents.
2026-06-21
update
OpenBMB/PilotDeck:Task-oriented AI Agent productivity platform
2026-06-21
update
earthtojake/text-to-cad:A collection of agent skills for CAD, robotics and hardware design
2026-06-21
update
microsoft/SkillOpt:SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
2026-06-21
update
op7418/guizang-ppt-skill:AI-agent Skill for generating polished HTML slide decks: editorial magazine and Swiss layouts, image prompts, social covers, and a WebGL/low-power presentation runtime.
2026-06-21
update
Slow breathing modulates brain function and risk behavior
2026-06-20
update
Introducing LifeSciBench
2026-06-20
update
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
2026-06-20
update
Effective Dimension Governs Generalization in Quantum Kernel Vision Models
2026-06-20
update
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-trainin
2026-06-20
update
Pitch Spelling Jazz Lead Sheets, Solo Transcriptions, Classical Piano and Monophonic Scores
2026-06-20
update
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact
2026-06-20
update
FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching
2026-06-20
update
CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia
2026-06-20
update
QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation
2026-06-20
update
A Multi-Agent system for Multi-Objective constrained optimization
2026-06-20
update
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
2026-06-20
update
Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference
2026-06-20
update
Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think
2026-06-20
update
The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse
2026-06-20
update
Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
2026-06-20
update
Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving
2026-06-20
update
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
2026-06-20
update
Statistical Properties of Training & Generalization
2026-06-20
update
Recurrent neural networks approximate continuous functions
2026-06-20
update
Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise
2026-06-20
update
Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining
2026-06-20
update
AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning
2026-06-20
update
CRAX: Fast Safe Reinforcement Learning Benchmarking
2026-06-20
update
Towards Modality-imbalanced Federated Graph Learning: A Data Synthesis-based Approach
2026-06-20
update
Pseudo-Feature Padding: A Lightweight Defense Against False Data Injection in Power Grids
2026-06-20
update
Sparsity, Superposition, and Forgetting: A Mechanistic Study of Representation Retention in Continual Learning
2026-06-20
update
SSH-Net: A Deep Neural Network for Predicting Failure Time Distribution Functions under Competing Risks with Application to GPU Data
2026-06-20
update
Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
2026-06-20
update
Optimal Order of Multi-Agent and General Many-Body Systems
2026-06-20
update
Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems
2026-06-20
update
Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
2026-06-20
update
LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
2026-06-20
update
Multi-Task Bayesian In-Context Learning
2026-06-20
update
Toward Calibrated Mixture-of-Experts Under Distribution Shift
2026-06-20
update
Optimal Deterministic Multicalibration and Omniprediction
2026-06-20
update
UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning
2026-06-20
update
How Transparent is DiffusionGemma?

关联动态

图卷积注意力:从谱视角看图去噪与扩散

本文从谱图理论出发,重新审视图卷积中的注意力机制,讨论其如何在图结构上同时实现去噪与扩散。作者将注意力权重解释为对不同频率成分的自适应滤波,说明其不仅影响邻居信息的聚合方式,也决定了信号在图上的平滑、保留边界与抑制噪声的平衡。文章进一步把该视角与图扩散过程联系起来,分析注意力在传播范围、稳定性和过平滑问题中的作用。该研究为理解图神经网络的表达能力提供了更统一的理论框架,也为设计更稳健的图学习模型提供了启发。

来源 arXiv 时间 2026-07-07 17:52:17
通过视觉动作结果推理对齐连接物理推理与任务泛化

本文聚焦具身智能中的一个核心难题:模型往往能在训练场景中完成动作,却难以将对物理世界的理解迁移到新任务与新环境。作者提出一种“视觉动作结果推理对齐”思路,把对动作后果的视觉推断与物理推理能力进行显式对齐,使模型不仅判断“该怎么做”,还学习“做完会发生什么”。这种机制有助于把局部感知、动作选择与结果预测统一起来,从而提升对未见任务的泛化能力。研究强调,真正稳健的任务执行不只依赖表面模式匹配,而需要可迁移的物理因果理解。

来源 arXiv 时间 2026-07-07 17:27:59
FreqDepthKV:面向长上下文推理的频率引导深度共享KV缓存压缩

这篇工作聚焦大模型长上下文推理中的KV Cache膨胀问题,提出 FreqDepthKV,一种结合“频率信息”与“深度共享”的缓存压缩方法。其核心思路是:并非所有注意力头或层对上下文记忆同等重要,可以利用表示频率特征来判断哪些KV更值得保留,并通过跨层共享降低冗余存储。相比传统只按重要性裁剪或简单量化的做法,FreqDepthKV更强调在压缩率与推理稳定性之间取得平衡,目标是在尽量少损失生成质量的前提下显著降低显存占用,提升长文本推理的可扩展性。

来源 arXiv 时间 2026-07-07 17:26:28
Pitwall:基于校准实时蒙特卡洛引擎的赛道策略自然语言简报

本文提出 Pitwall,一套将实时蒙特卡洛仿真与自然语言生成结合的赛车策略简报系统。它面向比赛中的战术决策场景,先通过校准后的概率模型持续评估安全车、进站、胎况与名次变化等关键变量,再把数值结果转写为面向工程师和车队的自然语言建议。与只输出分数或概率的传统系统不同,Pitwall 强调“忠实表达”——生成内容必须严格对应仿真结果,避免夸大、遗漏或不一致。研究重点在于把实时推演、置信度校准和可读性三者统一起来,从而为高时效、高风险的赛车策略提供可解释、可执行的语言化支持。

来源 arXiv 时间 2026-07-07 16:55:39
AirflowAttack:面向红外遥感视觉语言模型的热气流对抗扰动攻击

本文提出 AirflowAttack,一种针对红外遥感视觉语言模型的物理对抗攻击方法。与传统像素级扰动不同,该方法利用热气流在红外成像中的传播与扰动效应,通过构造可在真实环境中产生影响的热-气流扰动,干扰模型对遥感场景的视觉理解与语言对齐能力。研究表明,这类扰动不仅具有更强的环境可实现性,也能在红外遥感任务中显著降低模型的识别、描述与推理表现。该工作揭示了红外视觉语言模型在物理世界攻击面上的新脆弱性,并为后续开展鲁棒训练、传感器防护与对抗检测提供了参考。

来源 arXiv 时间 2026-07-07 16:46:46
真实世界数据复杂性下的大模型数据分析能力评测

这项研究聚焦大语言模型在真实数据分析场景中的表现,而不仅是标准化基准题。作者围绕现实世界数据常见的复杂性——如噪声、缺失值、格式不一致、异常点、混合类型字段以及多步骤推理需求——设计或整理了相应评测框架,用于检验模型在数据理解、清洗、分析与结论生成上的可靠性。研究强调,传统静态基准往往高估了模型能力,难以反映实际工作流中的脆弱性与误差累积。该工作有助于识别大模型在数据分析任务中的优势边界,也为构建更贴近生产环境的评测方法和智能数据工具提供参考。

来源 arXiv 时间 2026-07-07 16:43:05
量子多体系统时间演化预测的可证明学习分离

这篇论文研究如何从有限观测数据中学习量子多体系统的时间演化,并给出一个“可证明的学习分离”结果:对于某些时间演化预测任务,经典学习方法在样本或计算效率上会遭遇根本限制,而利用更合适的量子表征或假设则能够实现有效预测。作者围绕量子多体动力学中的可学习性,分析了模型表达、数据获取与泛化误差之间的关系,揭示了预测未来量子态演化并不只是工程问题,而是与系统结构和学习范式密切相关的理论问题。这一结论对量子模拟、量子控制以及量子机器学习中的动态建模具有参考价值。

来源 arXiv 时间 2026-07-07 16:30:09
WordVoice:面向大模型语音合成的词级多维显式解耦控制

这篇论文提出 WordVoice,一种面向基于大语言模型的文本转语音(TTS)系统的词级控制框架,重点解决语音生成中“想控制但控不准、多个属性互相干扰”的问题。作者将语音可控维度拆解为多个彼此解耦的词级属性,并通过显式建模让用户能够在词语粒度上分别控制语速、韵律、情感或强调等表现方式,而不是依赖粗粒度的整句指令。相比传统方法,WordVoice 更强调可解释性和组合式控制能力,适合需要精细口语化表达的场景,如有声内容制作、虚拟人播报和交互式语音助手。

来源 arXiv 时间 2026-07-07 16:22:59
从投票到智能体协作:面向BioASQ 14b的答案类型感知大模型流水线

本文聚焦 BioASQ 14b 生物医学问答任务,探讨如何依据问题的答案类型,构建更高效的大模型处理流水线。不同于传统依赖简单投票或统一提示词的做法,作者提出将问题按答案形式进行感知式路由,并进一步引入多智能体协作机制,让不同模型或代理分别承担检索、候选生成、证据整合与最终裁决等角色。该方法旨在提升复杂生物医学问题的覆盖率、准确性与可解释性,同时减少单一模型在长链推理和领域术语理解上的偏差。研究体现了从“多模型投票”向“协同式Agent系统”演进的趋势。

来源 arXiv 时间 2026-07-07 16:12:51
代理分析:作为条件编码器运行的视觉语言模型中的定位信号

本文研究视觉语言模型(VLM)在“条件编码器”模式下是否会保留可用于空间定位的内部信号。作者提出分析-by-proxy的视角:虽然模型并非显式训练来输出框或掩码,但在条件生成、理解或约束推理过程中,其表征可能已隐含对象位置与区域边界信息。通过探测不同层级与不同条件设置下的激活,研究揭示VLM中存在稳定的定位线索,并分析这些线索如何受提示词、视觉上下文和任务形式影响。该工作为理解多模态模型的空间感知能力、提升可解释性,以及将VLM用于弱监督定位与开放词汇视觉任务提供了新的依据。

来源 arXiv 时间 2026-07-07 16:11:13