Evaluating Fine-Tuning and Metrics for Neural Decompilation of Dart AOT Binaries
LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design
DynaKRAG: A Unified Framework for Learnable Evidence Control in Multi-Hop Retrieval-Augmented Generation
Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection
KVpop -- Key-Value Cache Compression with Predictive Online Pruning
Rethinking On-Policy Self-Distillation for Thinking Models
Weak-to-Strong Generalization via Direct On-Policy Distillation
A Retrieval-Augmented Framework for Detecting and Resolving Pragmatic Ambiguities in Natural Language Requirements
Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing
Wrong Before Right: Late Rescue and Interface Failure in Aligned Language Models
ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents
DecompRL: Solving Harder Problems by Learning Modular Code Generation
Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
Will Scaling Improve Social Simulation with LLMs?
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents
Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
CausalMix: Data Mixture as Causal Inference for Language Model Training
AGC-Bench: Measuring Artificial General Creativity
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
Adapting Foundation ASR Models to Dysarthric Speech: A Case Study
Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
Signed-Permutation Coordinate Transport for RMSNorm Transformers
Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding
Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering
Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index
Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
Experience Augmented Policy Optimization for LLM Reasoning
Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models
Know Before You Fetch: Calibrated Retrieval-Budget Allocation for Retrieval-Augmented Generation
When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding
FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models
Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing
Tandem Reinforcement Learning with Verifiable Rewards
RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning
Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
Reinforcement Learning without Ground-Truth Solutions can Improve LLMs
Steering Vision-Language Models with Joint Sparse Autoencoders
Black-Box Assisted Regression: Phase Transitions and Minimax Optimality
RAS: Measuring LLM Safety Through Refusal Alignment
WinDOM: Self-Family Distillation for Small-Model GUI Grounding
OpenThoughts-Agent: Data Recipes for Agentic Models
Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation
The Latent Bridge: A Continuous Slow-Fast Channel for Real-Time Game Agents
Qwen-AgentWorld: Language World Models for General Agents
Do LLM Embedding Spaces Recover Expert Structure?
GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation
POTracker: Optimizing Large Language Models for Standard-Compliant Power Outage Report Generation
Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior
Muown Implicitly Performs Angular Step-size Decay
Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability
BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories
Hypothesis-Driven Skill Optimization for LLM Agents
First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers
Sub-Billion, Super-Frontier: Small Language Models Rival Zero-Shot Frontier LLMs on General and Literary Relation Extraction
FleetAgent: Teleoperation Assistant for Autonomous Fleets via Vectorized V2N Messages
Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents
MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning
CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents
Dissecting Agentic RAG: A Component Ablation for Multi-Hop QA with a Local 7B Model
Good results fine tuning a local LLM like Qwen 3:0.6B to categorize questions
Two Qwen3 models on one DGX Spark: the residency math
Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families
SoftSkill: Behavioral Compression for Contextual Adaptation
Train, Retrieve, or Both? A Four-Arm Head-to-Head for Correct Statutory Citation on the Ontario Residential Tenancies Act
Judging to Improve: A De-biased VLM-as-3D-Judge Protocol for Single-Image 3D Generation
Probe-and-Refine Tuning of Repository Guidance for Coding Agents