LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#content-moderation

5 curated events
papersTODAY 04:00 UTC

Large-Scale Study Detects and Characterizes Ragebait on Japanese X

A new arXiv paper presents a method for identifying ragebait — content deliberately crafted to provoke outrage and boost engagement — at scale on the Japanese X platform. The authors move beyond binary detection toward characterizing how such posts are written and spread across a large sample. The work aims to fill gaps in reliable, large-scale measurement of ragebait online.

papersTODAY 04:00 UTC

SyRHM combines symbolic reasoning and associative retrieval for zero-shot harmful meme detection

Researchers propose SyRHM, a method for detecting harmful memes without task-specific training data. It targets implicit harm that comes from mismatches between image and text or from cultural stereotypes, which tripped up earlier multimodal detectors. The approach adds symbolic-language reasoning alongside associative retrieval to improve zero-shot performance.

papersSEP 10 04:00 UTC

Can foundation models moderate online content? Comparing instruction- and example-driven policies

A new arXiv paper investigates whether foundation models can apply complex content moderation policies reliably and consistently. The study compares two ways of translating moderation rules into model behavior: conveying them through explicit instructions versus through illustrative examples. The findings are relevant to platforms seeking scalable, automated moderation of online content.

papersSEP 12 04:00 UTC

Study Uses Bluesky's Public Moderation Logs to Map Harms and Automated Takedowns

A new arXiv paper examines Bluesky's content moderation system by analyzing its publicly accessible moderation logs, which most major platforms keep hidden. The authors characterize how much of the work is automated versus human-reviewed and catalog the types of harms that trigger moderation actions. The work argues that decentralized, transparent logging enables empirical moderation research that opaque platforms have long prevented.

papersSEP 12 04:00 UTC

SIRF: Spec-Internalized Risk Foundation Model for Industrial Content Risk Control

Researchers introduce SIRF, a foundation model designed for content risk control in industrial settings. The work argues that deployment success depends less on average accuracy and more on how much risky content can be automatically handled while maintaining high precision and sub-second response times. The model internalizes specification rules rather than relying on external filtering logic.