LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#vision-language-models

11 curated events
papersTODAY 04:00 UTC

arXiv paper examines robustness, cost and governance trade-offs in VLM document extraction

A new arXiv preprint argues that evaluations of vision-language models for extracting structured fields from business documents focus too heavily on accuracy against clean benchmarks. The authors propose assessing approaches along additional dimensions such as robustness, cost, and governance considerations, aiming to help practitioners pick a method suited to a given task complexity. No specific model or tool is released with the work.

papersTODAY 04:00 UTC

GroundBench benchmark aims to pinpoint where vision-language models fail on affordance tasks

A new arXiv paper introduces GroundBench, described as a factorized, counterfactual benchmark for identifying the specific points at which vision-language models break down on affordance tasks. The work cites a companion evaluation in which explicitly naming the target part in a manipulation prompt improved action accuracy by 0.32 to 0.63 across eight vision-language models, and no model exceeded a constant baseline before that part was named. The benchmark is intended to isolate these failures rather than report only aggregate scores.

papersTODAY 04:00 UTC

Robusto-2 Benchmark Tests Vision-Language Models for Self-Driving in Lima and New York

A new arXiv paper introduces Robusto-2, a benchmark evaluating both humans and vision-language models on autonomous driving tasks in Lima, Peru and New York City. The work targets how well multi-modal systems generalize when deployed in unfamiliar, out-of-distribution urban environments. It is a cross-listed replacement submission on arXiv cs.AI.

papersTODAY 04:00 UTC

VLM-Driven Parsing Method Aims to Improve Compositional SVG Generation

A new arXiv paper addresses the difficulty of producing structured, editable SVG graphics with vision-language models. Current approaches tend to output flat sets of paths with little semantic organization, which limits editability. The proposed method uses hierarchical semantic parsing driven by a VLM to make generation more compositional.

papersTODAY 04:00 UTC

Question-Guided Token Pruning Proposed as Privacy Defense for Vision-Language Models

A new arXiv paper proposes pruning visual tokens based on the question being asked, rather than sending all visual features to the server in split-learning setups. The approach targets vision-language question answering in federated, split, and U-shaped split learning, where raw data stays local but transmitted representations can still leak information. The authors frame selective transmission as a way to reduce both privacy exposure and bandwidth use.

papersTODAY 04:00 UTC

SportD benchmark tests whether vision-language models can make strategic soccer decisions

A new arXiv paper introduces SportD, a dataset that evaluates vision-language models on strategic decision-making using soccer as a testbed where action outcomes can be scored numerically. The work asks whether models that can describe a scene are also able to choose effective actions within it. This is a replacement version of the preprint.

papersTODAY 04:00 UTC

Paper measures and mitigates template collapse in 3D CT report generation

A new arXiv paper examines how 3D medical vision-language models can write fluent radiology-style reports while still missing critical findings and producing highly repetitive output. The authors characterize this behavior, which they call template collapse, noting that models default to generic phrasing that under-reports rare pathologies. They then propose ways to measure the problem and reduce it.

papersSEP 10 04:00 UTC

VANTAGE-Bench Measures the Infrastructure AI Gap in Vision-Language Models

A new arXiv paper introduces VANTAGE-Bench, a benchmark that tests how well vision-language models handle infrastructure-focused video as they move toward physical deployment. The authors argue that existing evaluations center on embodied AI using subject-centric consumer footage, leaving infrastructure AI largely unexamined. The benchmark is intended to quantify this gap between current model capabilities and real-world infrastructure monitoring needs.

papersSEP 10 04:00 UTC

RAU: Reference-based anatomical understanding for vision language models

Researchers introduce RAU, a reference-based approach that enables vision language models to identify, localize, and segment anatomical structures in medical images. By working from reference images instead of large volumes of expert annotations, the method addresses the shortage of labeled data that has slowed progress in medical image analysis.

papersSEP 12 04:00 UTC

arXiv paper proposes using vision-language models to automate classification error analysis

A new arXiv preprint describes a method that applies vision-language models to verification and validation of classification systems. The approach aims to replace the slow manual review of misclassified samples by automatically surfacing systematic failure patterns under realistic conditions. The authors frame the work as moving evaluation beyond standard benchmarks toward conditions closer to deployment.

papersSEP 12 04:00 UTC

TeleOCR Paper Targets Document Parsing for Both Digital and Camera-Captured Files

A new arXiv paper introduces TeleOCR, a method for converting unstructured documents into structured, machine-readable output. The work focuses on the gap between digitally born PDFs and photos of physical documents, a distinction that existing vision-language model approaches often handle unevenly. It falls within ongoing research into applying VLMs to document understanding tasks.