LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

low-resource-languages

topic17 events
modelsTODAY 04:00 UTC

MameLoshnLM: First Open-Source 8B Language Model for Yiddish Introduced

Researchers released MameLoshnLM, described as the first open-source 8-billion-parameter language model dedicated to Yiddish. The work also includes an evaluation benchmark intended to fill the gap in reliable testing resources for the language. It addresses the low digital availability of Yiddish text despite its substantial written heritage.

papersTODAY 04:00 UTC

Enemray: A Hassaniya-Centric Language Model Trained With a Stability-Plasticity Objective

Researchers present Enemray, a language model designed for general-purpose interaction in Hassaniya, a low-resource Arabic variety. Training follows a stability-plasticity objective intended to build strong linguistic and cultural competence while limiting degradation of the model's existing capabilities. The work is published as an arXiv preprint.

papersTODAY 04:00 UTC

Adaptive Language Sampling Method Targets Cross-Lingual Transfer for Low-Resource Languages

A new arXiv paper proposes an online adaptive sampling strategy for realigning multilingual language models, aiming to improve cross-lingual transfer to extremely low-resource languages. The authors note that existing realignment approaches typically use uniform or random sampling, which may underuse informative language pairs. Their method adjusts sampling dynamically as training proceeds within a distributed setup.

papersTODAY 04:00 UTC

Token Merging for Multilingual Speech Recognition Studied Across Model Scale

A new arXiv study systematically examines token merging as a way to cut the computational cost of large multilingual speech recognition models such as Whisper. The technique dynamically combines token representations during inference, and the authors test how its effectiveness varies with model size and fine-tuning. The work targets deployment efficiency for transcribing low-resource languages without language-specific training.

papersTODAY 04:00 UTC

arXiv paper proposes sequential adapter stacking for low-resource ASR

A new arXiv preprint describes a method that stacks adapters sequentially to help large multilingual speech recognition models handle low-resource languages. The authors note that current systems perform unevenly, favoring high-resource languages and losing accuracy where labeled audio is scarce. The abstract covers the approach at a high level, with results and evaluation details not included in the excerpt.

papersTODAY 04:00 UTC

FLORES+ Extended with Three Mozambican Bantu Language Evaluation Sets

Researchers added Portuguese-source evaluation data for Xichangana, Mozambican Nyanja, and Sena to the FLORES+ multilingual benchmark. The work examines how conflating closely related language varieties affects machine translation evaluation, comparing Xichangana against Tsonga and Mozambican Nyanja against Chichewa. It argues that merging distinct varieties into a single reference can distort measured translation quality.

papersTODAY 04:00 UTC

Bangla Sentence Function Classification Corpus and Benchmark Released

Researchers present a new annotated corpus for classifying sentence functions in Bangla, a resource previously lacking for the language. The work benchmarks several models on the task and adds interpretability analysis of their predictions. Such sentence-type identification supports dialogue systems, speech synthesis, and machine translation.

papersSEP 12 04:00 UTC

Switch-Aware Evaluation of ASR and Audio Language Models on English-Yoruba Code-Switched Speech

A new arXiv preprint argues that word error rate alone hides important failures when speech recognition systems and audio language models handle code-switched speech. The authors propose an evaluation method that accounts for language switches, and apply it to English-Yoruba audio, a low-resource pair with diacritics. They find that strong monolingual benchmark scores do not carry over to this setting.

papersSEP 12 04:00 UTC

Study Examines Multilingual LLM Weaknesses in Urdu and Low-Resource Languages

A new arXiv paper investigates how well multilingual large language models handle open-ended text generation in Urdu, a low-resource language. The authors argue that models marketed as multilingual often fall short in cultural and linguistic correctness outside high-resource languages. The work contributes to broader questions about the reliability of these systems for non-English users.

papersSEP 12 04:00 UTC

E-CONAN Benchmark Suite Targets Arabic Textual Entailment and Inference

A new arXiv paper introduces E-CONAN, a set of benchmarks covering entailment, contradiction and neutral relations for Arabic natural language inference. The authors frame the work as a response to the limited resources available for Arabic compared with English and other well-served languages, noting that inference models are a component of many downstream NLP applications. The datasets are intended to support training and evaluation of Arabic inference systems.

papersSEP 12 04:00 UTC

Study compares LLM adaptation methods for hate speech detection in Roman Urdu

A revised arXiv paper examines how large language models can be adapted to detect hate speech in Roman Urdu, a low-resource language written in Latin script. The authors compare several adaptation approaches, addressing challenges such as scarce annotated data, informal writing conventions, and the lack of standardized grammar. The work focuses on efficient methods suited to settings where labeled corpora are limited.

papersSEP 10 04:00 UTC

SALT Method Improves Token-Level Representations in Cross-Lingual Sentence Encoders

A new arXiv paper introduces SALT, a technique for strengthening how individual tokens are represented within multilingual sentence encoders. These encoders are optimized to align whole sentences across many languages, supporting applications like translation mining and zero-shot learning for low-resource languages. The paper targets the weaker token-level alignment that results from this sentence-focused training.

papersSEP 10 04:00 UTC

5-Dialects-BN Benchmark Probes How Transliteration Affects Bangla Dialect LLMs

A new arXiv paper introduces 5-Dialects-BN, a resource covering five Bangla dialects used to examine how transliteration choices influence large language model performance. The work targets the sharp accuracy drops LLMs show on low-resource, dialectally diverse languages, with Bangla being the world's sixth most spoken language. The authors position script and transliteration practices as an underexamined factor in this degradation.

papersSEP 10 04:00 UTC

BaltiVoice: 16.8-hour speech corpus and fine-tuned Whisper ASR system for Balti

Researchers have released BaltiVoice, a 16.8-hour read-speech corpus with 10,060 validated utterances for Balti, a Tibetic language spoken in Gilgit-Baltistan, Pakistan. The language previously had no publicly available speech recognition resources, making this the first open dataset and ASR model for Balti. The team fine-tuned OpenAI's Whisper architecture on the corpus to enable automatic speech recognition for the language.

papersSEP 10 04:00 UTC

Deterministic prompting for speaker-stable low-resource Greek TTS

Researchers present a data curation recipe combined with deterministic prompting to keep speaker identity stable in text-to-speech systems trained on limited speech data. Modern Greek serves as the test case, since it lacks the curated corpora that underpin state-of-the-art synthesis for high-resource languages. The work addresses quality degradation that TTS models typically show when clean training speech is scarce.

papersSEP 10 04:00 UTC

BuzzASR: 100+ monolingual fine-tuned Whisper models for speech recognition in 102 languages

A new arXiv paper introduces BuzzASR, a collection of more than one hundred monolingual Whisper models fine-tuned for automatic speech recognition in 102 languages. The language-specialized models are designed to cover languages that large general-purpose multilingual ASR systems often serve poorly. The release is presented as a resource for practitioners and researchers working on speech technology across many languages.

tipsAUG 28 00:00 UTC

Hugging Face Open ASR Leaderboard Adds First Global South Language

Hugging Face's Open ASR Leaderboard has expanded its coverage to include a language from the Global South for the first time. The addition broadens the benchmark's evaluation of automatic speech recognition systems beyond the predominantly high-resource languages it previously tracked. It reflects a wider push to measure model performance on underrepresented languages.

WHY IT MATTERS ↘Benchmarks drive where engineering effort goes, so extending a widely cited ASR leaderboard to a Global South language gives vendors and researchers a shared target for measuring quality on languages that commercial incentives alone have largely ignored. The caveat is that a single added language still reflects an underrepresentative sample, so teams should treat it as a starting signal for data collection and evaluation rather than evidence of broad multilingual coverage.