LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

DeepSeek

company10 events
tipsTODAY 05:30 UTC

Hands-on test compares Chinese AI models Kimi, Qwen, GLM and DeepSeek

A German tech outlet benchmarked four Chinese model families — Kimi, Qwen, GLM and DeepSeek — against established Western offerings such as ChatGPT and Claude. The review examines whether their strong benchmark scores translate into comparable quality in everyday practical use. It concludes that these alternatives offer notable capability at a lower price, though with trade-offs.

modelsSEP 12 05:56 UTC

DeepSeek releases v4.1-Flash, a 763B-parameter encoder-decoder model with vision

DeepSeek has introduced v4.1-Flash, a large-scale model with 763B total parameters built on a new causal encoder-decoder design that also handles vision input. Commentary from Latent Space and Sebastian Raschka argues the release is significant enough that it should have been branded as a new major version rather than a point update. Details on training, availability and licensing were not included in the report.

modelsSEP 12 02:46 UTC

DeepSeek v4.1 Flash runs at 23 seconds per token on 16GB M1 Mac Mini

A user report says the DeepSeek v4.1 Flash model can be loaded and run locally on a 2020 M1 Mac Mini with 16GB of unified memory, but generation is extremely slow at roughly 23 seconds per token. That pace makes the setup impractical for interactive use, though it shows the model can technically execute on older consumer hardware. The result highlights how limited RAM and memory bandwidth constrain local inference of large models.

papersSEP 11 04:00 UTC

HISA: Hierarchical Indexing Method for Fine-Grained Sparse Attention

A new arXiv paper proposes HISA, a hierarchical indexing approach for token-level sparse attention of the kind used in DeepSeek Sparse Attention. Such methods score every past key with a lightweight indexer before running attention on a selected subset, which becomes costly as context grows. HISA aims to make that key-selection step more efficient while keeping fine-grained selection quality.

modelsSEP 10 10:40 UTC

DeepSeek launches V4.1-Flash with 552B parameters and lower memory use

DeepSeek has released V4.1-Flash, a multimodal model with 552 billion total parameters that activates only 16 billion per token. The company says it cuts KV-cache memory to roughly a quarter of what its previous model required, which matters for agent workloads. It slightly edges out Opus 5 and GPT-5.6 Sol on the DeepSWE coding benchmark.

modelsSEP 10 06:11 UTC

DeepSeek publishes V4.1 Flash model on Hugging Face

DeepSeek has added a new model called V4.1 Flash to its Hugging Face repository, where the weights and model card are hosted. The listing drew attention on Hacker News, though the report gives no further detail on capabilities, size, or licensing. It appears to be a lighter or faster variant in the company's V4 series.

policySEP 9 20:46 UTC

NSA, CISA and FBI accuse six Chinese AI firms of model distillation

Three US agencies have named DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI as allegedly distilling American AI models at industrial scale since late 2024. The agencies also disputed DeepSeek's claim that its model was trained for about $5.6 million. They advised US providers to quietly reduce the quality of responses served to accounts flagged as suspicious.

policySEP 9 14:45 UTC

US agencies warn Chinese AI firms are extracting data from US models

US authorities say Chinese AI developers including DeepSeek and Alibaba are systematically pulling knowledge out of American models such as those built by OpenAI and Google. The agencies frame the activity as industrial espionage aimed at closing the capability gap at low cost. No details were given on which specific countermeasures might follow.

modelsSEP 9 11:19 UTC

DeepSeek plans V4.1 Flash release around September 10, 2026

DeepSeek intends to launch a new model called V4.1 Flash on or near September 10, 2026, Beijing time. According to the company, internal and external testing showed the Flash variant outperforming its V4 Pro model on performance, cost, speed, and task completion time. No independent benchmarks or pricing details were included in the announcement.