LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#peer-review

6 curated events
papersTODAY 04:00 UTC

Study Analyzes Self-Reported Limitations in NLP Research

A new arXiv paper examines the Limitations sections that top-tier NLP conferences have required since late 2022, noting that the volume of accepted papers has produced a corpus too large for manual review. The authors analyze these self-reported limitations to characterize what researchers themselves identify as the constraints of their work. The study aims to make this body of disclosures more tractable to assess at scale.

papersTODAY 04:00 UTC

AppliedScientist: Closed-Loop AI System Refines Papers via Iterative Reviewing

Researchers introduce AppliedScientist, a system that joins AI-generated peer review with automated revision in a closed loop. Instead of judging review quality only by how good the critique sounds, the work measures whether acting on that feedback actually improves the paper. The approach targets a gap in automated reviewing, where feedback is often evaluated in isolation from its effect on the manuscript.

papersTODAY 04:00 UTC

IntraGuard: Committee-Side Defenses Against Review Outsourcing to Chatbots

A new arXiv paper proposes IntraGuard, a defense meant to be deployed by editorial boards and program committees rather than by individual reviewers. It targets the growing problem of reviewers delegating peer review wholesale to commercial chatbots, a practice motivated by earlier evidence that these systems produce inadequate critiques. The work focuses on detecting such outsourced reviews from the committee's side.

papersYESTERDAY 12:54 UTC

Clay Institute says Navier-Stokes Millennium Prize problem apparently settled by AI

The Clay Mathematics Institute has indicated that the Navier-Stokes existence and smoothness problem, one of its seven $1 million Millennium Prize Problems, appears to have been resolved, reportedly with the help of AI. The institute says a formal verification process is now underway and stresses that its review procedure deliberately takes time. No prize determination has been made yet.

papersSEP 10 04:00 UTC

Study quantifies LLM-generated content in arXiv review vs non-review papers

A new arXiv paper examines how much AI-generated text appears in computer science review articles, after the platform stopped accepting unpublished reviews over concerns about machine-written submissions. The authors compare review and non-review papers to supply the quantitative evidence that was missing from the ban's rationale. The results may shape how academic platforms detect and regulate AI-written content.

papersSEP 12 04:00 UTC

NovGauge Benchmark Targets LLM Weakness in Judging Paper Novelty

A new arXiv preprint introduces NovGauge, a benchmark designed to test how well large language models assess the novelty of research papers. Unlike earlier benchmarks that reduce novelty to a single overall score, it breaks the task into separate dimensions so researchers can pinpoint where a model fails. The work is motivated by the growing use of LLMs in peer review at major AI conferences, where novelty judgments remain unreliable.