# AI News Flash — Daily Brief

**Date:** Friday, July 3, 2026  
**Type:** daily  
**Source:** https://ainewsflash.co/brief/48  
**Editors:** Justin Bunnell (https://www.linkedin.com/in/justinbunnell/), Laz Manrique (https://www.linkedin.com/in/laz-m-5a218b81/)  
**Publisher:** AI News Flash (https://ainewsflash.co)  
**License:** Republish with attribution.

---

## Key takeaways

- GPT-5.6 Sol runs at 750 tokens per second on Cerebras, unlocking real-time agents.
- xAI opens Grok Voice Agent Builder to all users, priced at $0.05 per minute.
- Claude Fable 5 returns after export suspension, leading SWE-Bench Pro at 80.3%.
- Senior SWE-Bench Opens a New Long-Horizon Coding Evaluation Tier
- Meituan open-sources LongCat-2.0, a 1.6T MoE trained entirely on Chinese ASICs.
- CISA's BOD 26-04 gives federal agencies three days to patch AI-exploitable critical flaws.

---

## Platforms

### GPT-5.6 Sol runs at 750 tokens per second on Cerebras, unlocking real-time agents.

OpenAI is deploying GPT-5.6 Sol on Cerebras hardware this month, targeting inference speeds of up to 750 tokens per second for enterprise customers. That figure is significant because it would make real-time agentic workflows practical at frontier-model quality, a threshold that slower inference has previously blocked. The partnership does not represent a broad rollout: Sol remains gated to approximately 20 government-approved partners, meaning the Cerebras channel serves a narrow slice of enterprises already cleared for access rather than the general enterprise market.

> Why it matters: Enterprises cleared for Sol access can now build low-latency agentic applications at frontier-model quality.

- Source: https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov

### xAI opens Grok Voice Agent Builder to all users, priced at $0.05 per minute.

xAI promoted its Grok Voice Agent Builder from developer beta to a generally available no-code platform on July 1, priced at $0.05 per minute for all users. The tool consolidates xAI's three-part voice API stack, covering speech-to-text, text-to-speech, and Grok Voice, into a single visual interface. Teams can now build and ship production voice agents without performing any API integration work, lowering the technical barrier to deploying conversational AI products on xAI infrastructure.

> Why it matters: Non-developer teams can now ship production voice agents on xAI infrastructure without writing API integration code.

- Source: https://www.basenor.com/blogs/news/xai-launches-grok-voice-agent-builder-beta-for-developers

## Capabilities

### Claude Fable 5 returns after export suspension, leading SWE-Bench Pro at 80.3%.

After a three-week export-control suspension, Anthropic restored Claude Fable 5 on July 1 across Claude.ai, Claude Code, Claude Platform, and Cowork. The model reclaims the top position on SWE-Bench Pro, a contamination-resistant benchmark using post-training-cutoff GitHub issues, scoring 80.3%, which is 11 points ahead of Opus 4.8 and more than 20 points ahead of GPT-5.5. Fable 5 also posts 95% on SWE-Bench Verified, 64.5% on Humanity's Last Exam with tools, and leads OSWorld-Verified for computer-use agents. Mythos 5, the same weights with safety classifiers removed, remains locked to Project Glasswing partners.

> Why it matters: Developers evaluating frontier coding agents now have a clear benchmark leader to test against across software engineering tasks.

- Source: https://www.datacamp.com/blog/claude-fable-5

### Senior SWE-Bench Opens a New Long-Horizon Coding Evaluation Tier

Factory AI launched Senior SWE-Bench this week as an open-source benchmark explicitly targeting vague, multi-step senior engineering tasks, the class of problems where current agents continue to struggle even as SWE-Bench Verified scores approach saturation above 90%. The benchmark is designed to surface gaps in long-horizon planning and ambiguous-specification resolution that single-file patching evaluations miss. No frontier model scores have been published yet, which means the benchmark currently serves as a forward-looking signal tool for tracking progress rather than a live ranking of existing models.

> Why it matters: AI builders and researchers now have a harder coding benchmark to expose planning weaknesses that existing saturated evaluations cannot detect.

- Source: https://www.theneurondaily.com/p/july-2-thursday

## Technology & Research

### Meituan open-sources LongCat-2.0, a 1.6T MoE trained entirely on Chinese ASICs.

Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter mixture-of-experts model that activates only 33 billion to 56 billion parameters per token and supports a native 1-million-token context window. It is the first frontier-scale model to complete full training and inference on a 50,000-card domestic Chinese ASIC cluster, likely Huawei Ascend 910C hardware. The model scores 59.5 on SWE-Bench Pro, edging GPT-5.5's 58.6, and 70.8 on Terminal-Bench, under an MIT license. Meituan employed custom 6D parallelism and binary-tree accumulation to compensate for a less mature chip software stack, demonstrating that frontier-scale training no longer requires Nvidia hardware.

> Why it matters: LongCat-2.0 demonstrates that frontier-scale model training is achievable on non-Nvidia hardware, reshaping assumptions about export controls and AI supply chains.

- Source: https://www.longcatai.org/models/longcat-2

## Regulation & Policy

### CISA's BOD 26-04 gives federal agencies three days to patch AI-exploitable critical flaws.

CISA's Binding Operational Directive 26-04 replaces two prior federal patching rules with a tiered risk-matrix model that gives civilian agencies three days to remediate the most dangerous vulnerabilities: those that are internet-facing, automatable, and capable of yielding full system control. Lower-severity issues may be deferred to the next upgrade cycle. The directive explicitly names AI-enabled exploit automation as a core prioritization factor, making it the first federal patching policy designed around adversaries who can weaponize vulnerabilities before patches are available. It fulfills a 30-day CISA mandate in the June 2 Trump AI Security Executive Order and supersedes KEV-catalog deadlines that had averaged 14 days in 2026.

> Why it matters: Federal agencies and their vendors must now remediate the most critical vulnerabilities within three days, compressing patching timelines significantly.

- Source: https://www.darkreading.com/cyber-risk/cisa-rewrites-federal-patching-requirements-ai-threat-era

### EU AI Act Omnibus is now law, shifting the high-risk compliance deadline to December 2026.

Following the Council of the EU's June 29 final adoption vote, the AI Act simplification package, known as the Digital Omnibus, was published in the EU Official Journal and entered into force on the third day after publication. The primary change shifts the high-risk AI system compliance deadline for most categories from August 2, 2026, to December 2, 2026. The Omnibus also introduces a new prohibition on AI-generated nudifier applications effective December 2, 2026, and extends watermarking transparency deadlines by four months. Companies that had been preparing for the August deadline are advised to retain it as a planning floor, since harmonized standards may not arrive until close to the new December dates.

> Why it matters: Enterprises deploying high-risk AI systems in the EU gain four additional months to achieve compliance, but planning buffers should remain in place.

- Source: https://knowledge.dlapiper.com/dlapiperknowledge/globalemploymentlatestdevelopments/2026/The-Digital-AI-Omnibus-Proposed-deferral-of-high-risk-AI-obligations-under-the-AI-Act


---

Archive: https://ainewsflash.co/archive  
RSS: https://ainewsflash.co/rss.xml  
LLM index: https://ainewsflash.co/llms.txt  
