LIVE PULSE
4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems1 src1.3 Paper proposes evolving context parameterization for large language models1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems1 src1.3 Paper proposes evolving context parameterization for large language models1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

video-understanding

topic3 events
papersTODAY 04:00 UTC

FriendBench Benchmark Tests Whether AI Can Tell Friends From Strangers

Researchers introduced FriendBench, a benchmark that evaluates how well humans and multimodal large language models can judge whether two people in a short video clip are already acquainted or meeting for the first time. The task uses 20-second recordings of ice-breaker conversations, where cues come from behavior and body language rather than spoken content alone. The work aims to measure social perception abilities that go beyond text-based reasoning.

papersSEP 10 04:00 UTC

Survey Reviews Inference-Efficiency Methods for Video and Audiovisual LLMs

A new survey on arXiv examines mechanisms for reducing inference costs in video large language models, which pair video representations with pretrained LLMs to generate responses from text prompts. The paper addresses why video understanding remains computationally expensive and organizes existing efficiency techniques across video and audiovisual tasks.

modelsSEP 1 17:08 UTC

Google DeepMind adds agentic video understanding to Gemini

Google DeepMind announced a new Gemini capability that lets the model analyze video content in an agentic, multi-step way rather than only answering single-pass questions about clips. The company says this allows the system to follow events over time, connect what it sees to tasks, and take further actions based on video input. Details on availability, pricing, and supported regions were not fully specified in the report.

WHY IT MATTERS ↘This shifts competition from benchmark video QA to deployable video agents that can chain perception with tools and actions, making continuous video analysis a practical automation layer for monitoring, editing, and interactive assistants. It also raises governance and cost questions, since always-on video ingestion and downstream actions increase privacy, liability, and compute demands that buyers will need to audit before production use.