通义千问 2.5

model Qwen2.5
AI 辅助整理 · 事实字段按公开来源校验

阿里巴巴通义千问开源系列,覆盖从小到大的多档位,多数尺寸以 Apache-2.0 开放,中文与代码能力强。

内容完整度: 5/5

事实字段

能力
通用/多语言
上下文
128K
参数量
0.5B–72B
开发方
阿里巴巴
许可证
Apache-2.0(部分)
来源追溯: 事实字段来自公开来源整理;字段级来源元数据待继续补齐,无法稳定核验的信息不做臆测。

整理说明

本页围绕「定位、事实字段、演进时间线、关联动态」组织信息,方便快速判断 通义千问 2.5 在 AI Agent 与大模型生态中的位置。

事实型字段保持保守:未公开或无法稳定核验的信息不做臆测;关联动态保留原始来源链接,便于继续阅读。

版本演进

2026-07-08
update
Evaluating Fine-Tuning and Metrics for Neural Decompilation of Dart AOT Binaries
2026-07-08
update
LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
2026-07-08
update
Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design
2026-07-08
update
DynaKRAG: A Unified Framework for Learnable Evidence Control in Multi-Hop Retrieval-Augmented Generation
2026-07-07
update
Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection
2026-07-07
update
KVpop -- Key-Value Cache Compression with Predictive Online Pruning
2026-07-07
update
Rethinking On-Policy Self-Distillation for Thinking Models
2026-07-07
update
Weak-to-Strong Generalization via Direct On-Policy Distillation
2026-07-07
update
A Retrieval-Augmented Framework for Detecting and Resolving Pragmatic Ambiguities in Natural Language Requirements
2026-07-07
update
Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
2026-07-07
update
Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing
2026-07-07
update
Wrong Before Right: Late Rescue and Interface Failure in Aligned Language Models
2026-07-07
update
ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents
2026-07-03
update
DecompRL: Solving Harder Problems by Learning Modular Code Generation
2026-07-03
update
Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
2026-07-03
update
Will Scaling Improve Social Simulation with LLMs?
2026-07-03
update
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
2026-07-03
update
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
2026-07-02
update
Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents
2026-07-02
update
Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
2026-07-02
update
CausalMix: Data Mixture as Causal Inference for Language Model Training
2026-07-02
update
AGC-Bench: Measuring Artificial General Creativity
2026-07-02
update
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
2026-07-01
update
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
2026-07-01
update
Adapting Foundation ASR Models to Dysarthric Speech: A Case Study
2026-07-01
update
Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
2026-07-01
update
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
2026-07-01
update
Signed-Permutation Coordinate Transport for RMSNorm Transformers
2026-07-01
update
Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?
2026-07-01
update
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding
2026-07-01
update
Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering
2026-07-01
update
Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index
2026-07-01
update
Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment
2026-06-30
update
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
2026-06-30
update
Experience Augmented Policy Optimization for LLM Reasoning
2026-06-30
update
Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models
2026-06-30
update
Know Before You Fetch: Calibrated Retrieval-Budget Allocation for Retrieval-Augmented Generation
2026-06-30
update
When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding
2026-06-29
update
FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models
2026-06-29
update
Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing
2026-06-29
update
Tandem Reinforcement Learning with Verifiable Rewards
2026-06-26
update
RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning
2026-06-26
update
Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA
2026-06-26
update
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
2026-06-26
update
Reinforcement Learning without Ground-Truth Solutions can Improve LLMs
2026-06-25
update
Steering Vision-Language Models with Joint Sparse Autoencoders
2026-06-25
update
BitNet Text Embeddings
2026-06-25
update
Black-Box Assisted Regression: Phase Transitions and Minimax Optimality
2026-06-25
update
RAS: Measuring LLM Safety Through Refusal Alignment
2026-06-25
update
WinDOM: Self-Family Distillation for Small-Model GUI Grounding
2026-06-24
update
OpenThoughts-Agent: Data Recipes for Agentic Models
2026-06-24
update
Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation
2026-06-24
update
The Latent Bridge: A Continuous Slow-Fast Channel for Real-Time Game Agents
2026-06-24
update
Qwen-AgentWorld: Language World Models for General Agents
2026-06-23
update
Do LLM Embedding Spaces Recover Expert Structure?
2026-06-23
update
GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation
2026-06-23
update
POTracker: Optimizing Large Language Models for Standard-Compliant Power Outage Report Generation
2026-06-23
update
Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior
2026-06-23
update
Muown Implicitly Performs Angular Step-size Decay
2026-06-23
update
Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability
2026-06-23
update
BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories
2026-06-23
update
Hypothesis-Driven Skill Optimization for LLM Agents
2026-06-23
update
First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers
2026-06-23
update
Sub-Billion, Super-Frontier: Small Language Models Rival Zero-Shot Frontier LLMs on General and Literary Relation Extraction
2026-06-23
update
FleetAgent: Teleoperation Assistant for Autonomous Fleets via Vectorized V2N Messages
2026-06-23
update
Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents
2026-06-23
update
MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning
2026-06-23
update
CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents
2026-06-23
update
Dissecting Agentic RAG: A Component Ablation for Multi-Hop QA with a Local 7B Model
2026-06-22
update
Good results fine tuning a local LLM like Qwen 3:0.6B to categorize questions
2026-06-21
update
Two Qwen3 models on one DGX Spark: the residency math
2026-06-20
update
Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families
2026-06-20
update
SoftSkill: Behavioral Compression for Contextual Adaptation
2026-06-20
update
Train, Retrieve, or Both? A Four-Arm Head-to-Head for Correct Statutory Citation on the Ontario Residential Tenancies Act
2026-06-20
update
Judging to Improve: A De-biased VLM-as-3D-Judge Protocol for Single-Image 3D Generation
2026-06-20
update
Probe-and-Refine Tuning of Repository Guidance for Coding Agents

关联动态

DynaKRAG:面向多跳检索增强生成的可学习证据控制统一框架

针对多跳检索增强生成中常见的证据噪声累积与路径次优问题,本文提出DynaKRAG框架。该工作突破传统启发式过滤局限,首创可学习的证据控制机制。通过统一的可微训练范式,模型能动态评估并加权多跳检索节点,精准剔除冗余干扰信息,显著优化复杂问答的生成准确率与事实一致性。该研究为构建高鲁棒性、可自主演进的大模型智能体知识底座提供了重要技术路径。

来源 arXiv 时间 2026-07-07 17:09:36
基于强化学习的LLM流程模型生成优化:奖励函数设计的核心作用

随着大语言模型在业务流程自动化与智能体工作流构建中的深入应用,其原生生成的流程模型常面临结构冗余与逻辑一致性不足的瓶颈。本研究引入强化学习范式,系统探究奖励函数设计对提升生成质量的关键作用。通过构建涵盖语法合规性、业务语义对齐及可执行性的多维度奖励信号,文章验证了精细化奖励机制在引导模型输出高保真流程结构中的核心效能。该研究为复杂工作流自动化建模、Agent任务编排及RLHF在垂直领域的工程落地提供了重要参考。

来源 arXiv 时间 2026-07-07 11:53:02
LongCrafter:基于证据图引导的指令合成与多样化长上下文理解框架

针对大模型长上下文理解中存在的推理单一与数据匮乏痛点,LongCrafter提出了一种基于证据图引导的指令自动合成新范式。该方法通过构建多源事实证据的逻辑图谱,自动化生成高多样性、强逻辑关联的长文本训练指令,有效缓解模型在长文档场景下的注意力稀释与幻觉问题。实验表明,该框架能显著提升模型在复杂长程依赖任务中的表现,为构建高质量长上下文对齐数据提供了可扩展的新路径。

来源 arXiv 时间 2026-07-07 11:35:48
面向Dart AOT二进制的神经反编译:微调策略与评估指标研究

随着大模型在代码智能领域的渗透,传统反编译技术正加速向神经网络范式演进。本文聚焦Flutter生态核心的Dart AOT二进制文件,系统评估了不同大模型微调策略在神经反编译任务中的表现。针对AOT编译带来的控制流扁平化与符号剥离难题,研究构建了多维度评估指标体系,不仅关注语法还原度,更深入验证语义等价性与功能可执行性。该工作为移动端应用安全审计、恶意代码分析及跨平台逆向工程提供了可靠的基线模型与量化标准,对推动AI驱动的二进制代码理解具有重要参考价值。

来源 arXiv 时间 2026-07-07 10:36:12
弱模型驱动强模型泛化:直接同策略蒸馏新范式

针对大模型高质量标注数据稀缺的瓶颈,本文提出基于直接同策略蒸馏的弱至强泛化框架。该方法突破传统离线教师监督局限,直接在强模型自采样轨迹上进行策略蒸馏,通过直接优化算法将弱监督信号高效对齐至强表征。实验证实,该机制能显著突破弱监督性能天花板,为降低大模型对齐成本与探索高效数据利用范式提供重要参考。

来源 arXiv 时间 2026-07-06 17:59:58
思维模型策略内自蒸馏机制的重新审视与优化路径

本文系统探讨大语言模型在复杂推理任务中采用的策略内自蒸馏技术,揭示传统方法在知识迁移效率与推理稳定性方面的局限。通过构建动态反馈优化框架,提出分层知识蒸馏与策略对齐新范式,在数学推理、代码生成等任务中验证了模型思维链的连贯性提升。研究指出当前自蒸馏过程易陷入策略漂移与表征坍缩,并给出基于对抗正则化与多阶段微调的改进方案,为下一代可解释推理模型提供理论支撑与工程实践参考。

来源 arXiv 时间 2026-07-06 15:01:35
KVpop:基于预测性在线剪枝的大模型键值缓存压缩

在LLM推理中,KV缓存随序列线性增长,成为长上下文与高吞吐场景的核心瓶颈。传统压缩方法多依赖事后评估,难以兼顾实时调度与模型精度。本文提出KVpop,一种预测性在线KV缓存压缩框架。该方法引入轻量级预测模块,在自回归生成阶段动态预判注意力重要性,实时剪枝冗余键值对,显著降低显存占用与访存延迟。基准测试显示,该方案在大幅压缩缓存的同时维持了生成质量,为移动端部署与超长文本处理提供了高效的推理优化路径。

来源 arXiv 时间 2026-07-06 13:32:34
突破独立标签局限:基于施瓦茨几何的人类价值观解码

传统价值观检测常将各类价值视为孤立标签,忽略了人类价值体系内在的关联与冲突。本文提出基于施瓦茨价值圆环几何结构的解码框架,通过建模价值观间的正交、相邻与对立关系,实现对文本中复杂价值取向的精准推断。该方法将离散分类转化为连续几何空间中的结构化推理,显著提升大模型在伦理对齐、智能体价值推理与跨文化分析中的可解释性,为构建具备人类价值感知能力的AI生态提供新范式。

来源 arXiv 时间 2026-07-06 13:26:01
ToolFailBench:大模型智能体工具调用失效诊断基准

随着大模型智能体广泛依赖外部工具执行复杂任务,工具调用失效已成为制约其可靠性的核心瓶颈。现有评测多聚焦整体成功率,缺乏对失败根因的细粒度归因。本文提出ToolFailBench,首个专注智能体工具调用失效的诊断基准。该框架系统划分了意图偏差、参数构造错误、执行流断裂与异常恢复缺失等故障类别,并配套自动化诊断流水线与多维评估指标。研究者可借此精准定位工具链薄弱环节,针对性优化规划逻辑与容错机制,为打造高鲁棒性Agent生态提供标准化基础设施。

来源 arXiv 时间 2026-07-06 05:25:21
错在先对在后:对齐大模型的延迟救援与接口失效

本文揭示了对齐大语言模型在自回归生成中的时序性缺陷:模型常在初期产生错误或违规内容,而安全对齐机制往往在序列中后期才触发“延迟救援”,导致底层生成逻辑与外部安全接口发生断裂。研究指出,现有对齐范式缺乏细粒度时序控制,易引发Agent行为失控。该发现为重构实时拦截架构、提升人机交互鲁棒性提供了关键理论支撑。

来源 arXiv 时间 2026-07-06 03:51:24

相关实体