LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#distillation

23 curated events
papersTODAY 04:00 UTC

Joint-Output On-Policy Distillation Targets Output-Mode Gap in Speech Language Models

A new arXiv paper addresses a mismatch that arises when speech language models autoregressively generate interleaved text and acoustic tokens. The authors propose a joint-output on-policy distillation approach intended to close this output-mode gap while preserving the streaming and text-guided benefits of the interleaved design. The work falls under computation and language research and has not yet been peer reviewed.

papersTODAY 04:00 UTC

arXiv paper uses persona-based distillation to improve LLM humor generation

A new arXiv preprint argues that standard next-token training objectives work against the surprise and incongruity that comedy requires, making humor a hard task for large language models. The authors propose HumorGen, a method that distills knowledge from multiple personas to create a cognitive synergy effect. The work is a revised submission and focuses on generation quality rather than a released product.

papersTODAY 04:00 UTC

Coupled-Noise Distillation Method Targets One-Step Block Generation in Diffusion Language Models

A revised arXiv paper examines why masked diffusion language models can produce incoherent text blocks: they decode every position in a block in parallel from separate marginal distributions. The authors propose a coupled-noise distillation approach intended to let such models generate a whole block in a single step while keeping the tokens mutually consistent. The work is a research preprint and has not been peer-reviewed.

papersTODAY 04:00 UTC

Study Examines Training Domain Specialists Without Reasoning Trajectories

A new arXiv paper looks at domain expert distillation, where a teacher model normally passes reasoning traces to a student model. It investigates what happens when specialists are trained only on question-answer pairs and no explicit reasoning supervision is provided. The work aims to clarify how much the reasoning trajectories actually contribute to the student's performance.

papersSEP 10 04:00 UTC

Paper examines distilling synthetic data for time series foundation models

A new arXiv preprint looks at how time series foundation models are pretrained on artificially generated trajectories where the underlying data-generating process is known. The work focuses on distillation methods rather than the conventional loss-based pretraining objectives that compare model outputs against targets. It aims to improve how these models learn from synthetic time series data.

papersTODAY 04:00 UTC

arXiv Paper Probes Dataset Biases Behind Phantom Transfer

A new preprint on arXiv studies why a teacher model's bias can still pass to a student model even when the training data has had all overt mentions of that bias removed. The authors report that no data-level defense tested so far reliably detects or eliminates this residual, or "phantom," transfer. The work frames the phenomenon as a dataset-level problem rooted in subtle statistical traces rather than explicit labels.

papersTODAY 04:00 UTC

arXiv Paper Examines How Much Training Data Matters in On-policy Distillation

A new arXiv preprint investigates how much of the benefit from on-policy distillation actually comes from the training data used. Testing the two teacher-student pairings most often seen in practice, the authors report findings that challenge assumptions about data's role in the method. The work is positioned as a closer look at a technique that has become standard in frontier post-training pipelines.

papersTODAY 04:00 UTC

Verifier-Gated Multi-Expert Distillation Aimed at Scientific Reasoning

A new arXiv paper examines multi-teacher on-policy distillation, the technique of training specialist models and then transferring their abilities to a single student using the student's own generated outputs. The authors propose assigning supervision token by token rather than sequence by sequence, with a verifier deciding which expert teacher should guide each token. The method is aimed at scientific reasoning tasks.

papersTODAY 04:00 UTC

Stopping and restarting strategy speeds up multi-turn agentic on-policy distillation

A new arXiv paper addresses the high cost of on-policy distillation, which relies on expensive autoregressive rollouts by the student model and scales poorly when tasks span multiple turns. The authors propose deciding when to halt a rollout and where to resume it, aiming to cut the compute spent on generating student trajectories. The method targets more efficient transfer of capabilities from large teacher models to smaller students in agentic settings.

papersTODAY 04:00 UTC

Counteraction-Aware Multi-Teacher Distillation Aims to Preserve LLM General Skills

A new arXiv paper addresses how domain-specific fine-tuning can erode the general abilities an LLM originally had. The authors propose a counteraction-aware extension of multi-teacher on-policy distillation, which trains on student-generated text under multiple teachers to restore lost capabilities while keeping domain performance. The method targets the trade-off between specialization and retaining broad competence.

papersTODAY 04:00 UTC

REGEN paper proposes replay-recycling for expert-to-generalist LLM distillation via offline RL

A revised arXiv paper introduces REGEN, a method that recycles replay data to distill specialized expert policies into a more general model using offline reinforcement learning. The approach targets the cost of scaling online RL, which is widely used to develop long-horizon reasoning and tool-use abilities in large language models. The v3 revision appears in both the cs.AI and cs.LG listings.

papersTODAY 04:00 UTC

Qwen-Image-Flash: Rethinking the Training Recipe for Few-Step Distillation

A new arXiv paper presents Qwen-Image-Flash, a revised training approach for few-step distillation in visual generative models. The method aims to cut inference costs and enable real-time image generation without sacrificing output quality. It targets efficient deployment of generative foundation models in production settings.

papersTODAY 04:00 UTC

Temporal Self-Distillation Speeds Up Discrete Diffusion Language Models

A new arXiv paper proposes Temporal Self-Distillation, a training method aimed at discrete diffusion language models that generate several tokens at once. Such models lose quality when too many tokens are decoded in parallel, and the technique is presented as a simple way to reduce that degradation. The approach targets faster inference without the accuracy drop that usually accompanies aggressive parallel decoding.

papersTODAY 04:00 UTC

Attention Bridge Method Distills Transformers into Mamba Models with Less Data

A new arXiv paper proposes an "attention bridge" technique for converting pretrained Transformer models into Mamba-style state-space models. The approach aims to make the distillation process more data efficient, addressing the high compute cost of training competitive SSMs from scratch. The work targets the gap between the mature Transformer ecosystem and the less developed tooling around state-space architectures.

papersSEP 10 04:00 UTC

CompassOPD adapts on-policy distillation to cross-family model pairs

New research introduces CompassOPD, a method that extends on-policy distillation to settings where the teacher and student models come from different families. It derives within-family likelihood shifts to provide dense token-level supervision on student-generated outputs, tackling the effectiveness drop standard OPD exhibits in cross-family scenarios.

papersSEP 10 04:00 UTC

RouteBridge paper introduces bidirectional distillation between NeRF and 3D Gaussian Splatting

A new arXiv preprint presents RouteBridge, a framework for transferring knowledge between neural radiance fields and 3D Gaussian Splatting, two 3D scene representations with complementary strengths. Rather than fixing one representation as the teacher for an entire scene, the method routes distillation bidirectionally, relying on the more reliable representation for each part of the scene.

papersSEP 10 04:00 UTC

On-Policy Distillation Proposed for Vision-Language Model Adaptation on Low-Quality Data

A new arXiv paper introduces an on-policy distillation approach for adapting compact vision-language models from a larger task-trained teacher. Rather than relying solely on teacher predictions as training targets, the method lets the student learn from its own outputs, which the authors report makes it especially effective when multimodal training data is noisy or low quality.

papersSEP 10 04:00 UTC

arXiv paper proposes data-centric post-training pipeline for financial reasoning

A new research paper tackles the shortage of training data suitable for reasoning-focused fine-tuning in the financial domain, noting that most available QA pairs lack explicit reasoning steps, sufficient context, or reliably checkable answers. The authors present a pipeline that mines financial text, distills it into reasoning-oriented training examples, and applies learning with verifiable answers to improve model performance on financial tasks.

papersSEP 10 04:00 UTC

Zone of Proximal Policy Optimization: teacher guidance via prompts, not gradients

A new arXiv paper argues that knowledge distillation breaks down when the student model is much smaller than its teacher, because imitating the teacher's logits locks the student into its sharpest output modes and harms generalization. The authors propose letting the large teacher guide the small student through prompts during reinforcement-learning fine-tuning instead of through gradient-based distillation. The work appears in the computational linguistics category on arXiv.

papersSEP 10 04:00 UTC

Study Examines Decision Shifts and Grounding in Correctness-Gated Multi-Teacher Distillation

A new arXiv paper in cs.AI reports a controlled comparison of eight correctness-gated multi-teacher distillation setups, all sharing the same data sources and a 63.9-million-parameter student model. The authors argue that achieving correct candidate decisions is a different objective from keeping model rationales grounded, and their audit of grounding ended inconclusive. The experiments also documented shifts in decisions and a loss of label functionality across the tested configurations.

papersSEP 12 04:00 UTC

X-AuT Compresses Speech LLM Audio Encoders via Cross-Scale Distillation

Researchers propose X-AuT, a method that progressively compresses the audio encoder of speech large language models rather than deleting whole blocks at once. Because abrupt block removal distorts the embeddings the decoder receives and leads to word deletion and premature end-of-sequence errors, the approach uses cross-scale distillation to shrink encoder depth while preserving output quality. The aim is to cut inference cost without the accuracy loss typical of standard pruning.

papersSEP 12 04:00 UTC

arXiv paper proposes unified per-token gating family for on-policy distillation

A new arXiv preprint introduces a family of per-token gating methods for on-policy knowledge distillation that mixes forward and reverse KL losses. The authors argue that prior approaches such as EOPD and ToDi each rely on a single fixed gating signal, and their framework generalizes these with multi-channel and bias coefficients. The work is a methodological contribution aimed at improving how distillation losses are weighted per token during training.

policySEP 9 20:46 UTC

NSA, CISA and FBI accuse six Chinese AI firms of model distillation

Three US agencies have named DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI as allegedly distilling American AI models at industrial scale since late 2024. The agencies also disputed DeepSeek's claim that its model was trained for about $5.6 million. They advised US providers to quietly reduce the quality of responses served to accounts flagged as suspicious.