LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#low-resource

15 curated events
papersTODAY 04:00 UTC

arXiv paper proposes sequential adapter stacking for low-resource ASR

A new arXiv preprint describes a method that stacks adapters sequentially to help large multilingual speech recognition models handle low-resource languages. The authors note that current systems perform unevenly, favoring high-resource languages and losing accuracy where labeled audio is scarce. The abstract covers the approach at a high level, with results and evaluation details not included in the excerpt.

papersTODAY 04:00 UTC

Mizan benchmark evaluates LLMs on Iraqi Arabic and civic context

Researchers introduced Mizan, a benchmark designed to test large language models on Iraqi Arabic and on civic topics relevant to Iraq. Existing Arabic evaluation efforts have largely centered on Modern Standard Arabic, leaving regional dialects and country-specific knowledge thinly covered. The work aims to give a national-level measure of model performance beyond aggregated MSA leaderboards.

papersTODAY 04:00 UTC

Adaptive Language Sampling Method Targets Cross-Lingual Transfer for Low-Resource Languages

A new arXiv paper proposes an online adaptive sampling strategy for realigning multilingual language models, aiming to improve cross-lingual transfer to extremely low-resource languages. The authors note that existing realignment approaches typically use uniform or random sampling, which may underuse informative language pairs. Their method adjusts sampling dynamically as training proceeds within a distributed setup.

papersTODAY 04:00 UTC

FLORES+ Extended with Three Mozambican Bantu Language Evaluation Sets

Researchers added Portuguese-source evaluation data for Xichangana, Mozambican Nyanja, and Sena to the FLORES+ multilingual benchmark. The work examines how conflating closely related language varieties affects machine translation evaluation, comparing Xichangana against Tsonga and Mozambican Nyanja against Chichewa. It argues that merging distinct varieties into a single reference can distort measured translation quality.

papersTODAY 04:00 UTC

Melbourne WMT 2026 Submission Targets Pacific Creole Translation

Researchers from the University of Melbourne submitted a system to the WMT 2026 Creole Language Translation shared task, covering Tok Pisin, Bislama, and Solomon Pijin. The work emphasizes balanced performance across a wide range of domains rather than a single text type. It relies on pre-training followed by domain-aware fine-tuning for these low-resource Pacific creoles.

papersTODAY 04:00 UTC

Enemray: A Hassaniya-Centric Language Model Trained With a Stability-Plasticity Objective

Researchers present Enemray, a language model designed for general-purpose interaction in Hassaniya, a low-resource Arabic variety. Training follows a stability-plasticity objective intended to build strong linguistic and cultural competence while limiting degradation of the model's existing capabilities. The work is published as an arXiv preprint.

papersTODAY 04:00 UTC

Study tests reference-free triage of LLM translation errors in Pali texts

A new arXiv paper examines how to decide which machine translations of classical texts need expert review when no human reference translation exists. Using Pali-to-English as a test case, it combines source-novelty measures, GEMBA-style quality scoring, and a limited review budget. The goal is a practical way to route scarce expert attention to the most error-prone outputs.

papersTODAY 04:00 UTC

ParsVoice: Large Multi-Speaker Persian Speech Corpus for Text-to-Speech

Researchers released ParsVoice, described as the largest publicly available multi-speaker Persian speech corpus, aimed at addressing the scarcity of open speech-text data for the language. The dataset is intended to support work on multi-speaker text-to-speech, speech-language modeling, and other low-resource speech tasks. The paper was posted to arXiv under computer science categories.

papersSEP 10 04:00 UTC

5-Dialects-BN Benchmark Probes How Transliteration Affects Bangla Dialect LLMs

A new arXiv paper introduces 5-Dialects-BN, a resource covering five Bangla dialects used to examine how transliteration choices influence large language model performance. The work targets the sharp accuracy drops LLMs show on low-resource, dialectally diverse languages, with Bangla being the world's sixth most spoken language. The authors position script and transliteration practices as an underexamined factor in this degradation.

papersSEP 10 04:00 UTC

Deterministic prompting for speaker-stable low-resource Greek TTS

Researchers present a data curation recipe combined with deterministic prompting to keep speaker identity stable in text-to-speech systems trained on limited speech data. Modern Greek serves as the test case, since it lacks the curated corpora that underpin state-of-the-art synthesis for high-resource languages. The work addresses quality degradation that TTS models typically show when clean training speech is scarce.

papersSEP 10 04:00 UTC

Study proposes evaluation of diachronic semantic change in Sinhala

A new arXiv paper examines how word meanings in Sinhala have shifted over long historical periods, focusing on a language with limited textual resources. The authors tackle obstacles such as sparse historical corpora and the drawbacks of static embedding alignment techniques when measuring semantic drift. The work contributes an evaluation approach for diachronic semantic change in low-resource NLP settings.

papersSEP 10 04:00 UTC

Stress-Aware Sentence-Level Filipino G2P Model Built With Weakly-Supervised ByT5 Fine-Tuning

A new arXiv preprint presents a fine-tuned ByT5 model that converts Filipino text into phoneme sequences at the sentence level while also predicting stress. Since Filipino's spelling maps to sounds fairly predictably, the authors focus on the harder problem of capturing prosody, which they tackle using weakly supervised training data.

papersSEP 12 04:00 UTC

Study compares LLM adaptation methods for hate speech detection in Roman Urdu

A revised arXiv paper examines how large language models can be adapted to detect hate speech in Roman Urdu, a low-resource language written in Latin script. The authors compare several adaptation approaches, addressing challenges such as scarce annotated data, informal writing conventions, and the lack of standardized grammar. The work focuses on efficient methods suited to settings where labeled corpora are limited.

papersSEP 12 04:00 UTC

Study Examines Multilingual LLM Weaknesses in Urdu and Low-Resource Languages

A new arXiv paper investigates how well multilingual large language models handle open-ended text generation in Urdu, a low-resource language. The authors argue that models marketed as multilingual often fall short in cultural and linguistic correctness outside high-resource languages. The work contributes to broader questions about the reliability of these systems for non-English users.

papersSEP 12 04:00 UTC

Switch-Aware Evaluation of ASR and Audio Language Models on English-Yoruba Code-Switched Speech

A new arXiv preprint argues that word error rate alone hides important failures when speech recognition systems and audio language models handle code-switched speech. The authors propose an evaluation method that accounts for language switches, and apply it to English-Yoruba audio, a low-resource pair with diacritics. They find that strong monolingual benchmark scores do not carry over to this setting.