5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.3 Study traces LLM hallucinations to competing latent associations — 1 src1.3 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text — 1 src1.3 Study Analyzes Self-Reported Limitations in NLP Research — 1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.3 Study traces LLM hallucinations to competing latent associations — 1 src1.3 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text — 1 src1.3 Study Analyzes Self-Reported Limitations in NLP Research — 1 src
Study quantifies when multi-agent LLM coding is reliable for qualitative analysis
Researchers examined how AI coding agents collaborate, disagree, and converge when performing multi-coder qualitative coding, a task where their usefulness has been assumed but rarely measured. The paper reports empirical results on the settings and conditions under which multi-agent LLM coding performs dependably, and highlights the gaps that limit its reliability. It frames these findings as both challenges and opportunities for building better LLM-assisted qualitative research tools.