AI News Flash · Daily Brief

GPT-5.6 is now the default brain inside Microsoft 365 Copilot for enterprise users.

Platforms

GPT-5.6 is now the default brain inside Microsoft 365 Copilot for enterprise users.

OpenAI announced on July 14 that GPT-5.6 is now the default model powering Microsoft 365 Copilot, replacing earlier GPT-5.x versions across Word, Excel, PowerPoint, Chat, and Cowork. The upgrade puts OpenAI's latest flagship directly inside the productivity suite used by hundreds of millions of enterprise workers. Sam Altman has stated that GPT-5.6 is 54% more token-efficient on agentic coding tasks than its predecessor, a gain that could meaningfully reduce inference costs for enterprises running high-volume automated workflows through Copilot. The move also deepens the operational dependency between Microsoft and OpenAI at the model level.

Why it matters: Enterprise IT teams must now reassess Copilot workflow performance and cost assumptions based on GPT-5.6 behavior.

Anthropic gives US K-12 teachers free premium Claude access backed by FERPA protections.

Anthropic launched Claude for Teachers, a program offering verified US K-12 educators free access to premium Claude capabilities, a curated library of teaching skills, and curricula mapped to academic standards across all 50 states. The initiative is backed by a partnership with the American Federation of Teachers and includes a K-12 data processing addendum written to FERPA compliance, with student data explicitly excluded from model training. By entering the education market ahead of back-to-school AI adoption decisions, Anthropic positions Claude as a trusted institutional tool and competes for the same educator mindshare that Google and Microsoft have cultivated through their own education programs.

Why it matters: School districts evaluating AI tools now have a FERPA-compliant Claude option that could shape educator and student AI habits long term.

Capabilities

Unisound U2 hits frontier reasoning scores at $0.15 per million tokens.

Unisound's U2, a mixture-of-experts model with 266 billion total parameters but only 10 billion active per token, has posted independently verified scores of 86.9% on GPQA Diamond, 85.8% on MATH-500, and 72.2% on SWE-bench Verified. Its pricing of $0.15 per million input tokens and $0.30 per million output tokens places frontier-grade reasoning performance at a cost roughly an order of magnitude below comparable closed models. The 10B active parameter footprint keeps inference costs low even on complex reasoning tasks that previously demanded far larger active parameter counts to reach this score range on GPQA Diamond, making high-end reasoning accessible to teams with limited inference budgets.

Why it matters: AI teams running cost-sensitive reasoning workloads now have a verified low-cost alternative to closed frontier models worth serious evaluation.

Google DeepMind targets Gemini 3.5 Pro launch after scrapping and rebuilding the model.

Google DeepMind is targeting July 15 for the general availability of Gemini 3.5 Pro on Vertex AI, but the launch follows a full model rebuild after engineers discovered structural failures in recursive tool-calling and SVG generation in the original version. The rebuilt model is reported to feature a 2-million-token context window and support for up to 8 concurrent function calls per turn. As of July 13, no model card, pricing page, or public API listing had appeared, leaving all specifications, including the launch date itself, unconfirmed. If the date holds, Gemini 3.5 Pro enters a market where Claude Fable 5 and GPT-5.6 are already in active production.

Why it matters: Enterprises evaluating long-context models should wait for confirmed Gemini 3.5 Pro specs before committing to integration plans.

Technology & Research

DeepSeek-V4 cuts KV cache 90% at 1M-token context with two new attention methods.

DeepSeek-AI's technical report, arXiv 2606.19348, details a hybrid attention architecture pairing Compressed Sparse Attention and Heavily Compressed Attention. CSA consolidates every m KV-cache tokens into a single entry and applies top-k selection, while HCA applies heavier compression with dense attention over the result. At one-million-token context, V4-Pro runs at 27% of V3.2's single-token inference FLOPs and just 10% of its KV-cache footprint, with no accuracy regression on reported benchmarks. DeepSeek releases two open-weight variants: a 1.6T-total, 49B-active Pro and a 284B-total, 13B-active Flash, both under MIT license, putting this inference-cost reduction within reach of any team operating a multi-GPU cluster.

Why it matters: Developers building long-context applications can now access a credible open-weight architecture that slashes memory and compute costs at scale.

Moonshot Kimi K2.7 Code beats K2.6 by 21.8% on coding benchmarks in just six weeks.

Moonshot AI released Kimi K2.7 Code on June 12, a coding-specialized open-weight model developed through continual pretraining on the K2.6 base, which carries 1 trillion total parameters and 32 billion active per token. Released under a Modified MIT license, K2.7 Code scores 21.8% higher than K2.6 on Kimi Code Bench v2 and tops open-weight leaderboards for long-horizon coding agents as of July 2026. The six-week gap between major coding updates signals that open-weight labs are now iterating on coding specialization at a pace that rivals or exceeds the release cadence of closed commercial model providers.

Why it matters: Open-weight coding model iteration is now fast enough that AI engineering teams should reassess their model choices on a monthly basis.

Regulation & Policy

The NO FAKES Act passes Senate Judiciary and moves toward a full Senate vote.

The NO FAKES Act, S. 4591, passed the Senate Judiciary Committee unanimously and now heads to a full Senate floor vote, making it the first federal bill to establish a licensable property right in individual voice and visual likeness, with civil liability attached to unauthorized AI-generated replicas. Lead sponsor Sen. Marsha Blackburn is negotiating with the White House to fold the bill into a broader AI preemption package alongside the Kids Online Safety Act and age-verification requirements. The bill preempts future state digital-replica laws while preserving existing ones, placing it at the center of the ongoing federal-versus-state AI governance debate affecting every company that generates synthetic media.

Why it matters: Companies building or deploying AI voice and likeness tools must prepare for potential federal civil liability if the NO FAKES Act clears the Senate.

NIST secures pre-deployment AI model access from Google DeepMind, Microsoft, and xAI.

NIST's Center for AI Standards and Innovation has signed agreements with Google DeepMind, Microsoft, and xAI that give the agency access to their frontier models for pre-deployment evaluations and targeted security research. These are described as the first formal government access arrangements of this kind with frontier AI developers, moving the NIST AI Risk Management Framework from voluntary guidance toward structured pre-market assessment. The approach parallels the EU AI Act's cybersecurity action plan, which is pursuing a similar third-party evaluation capacity through a planned open call. The deals create a template that other frontier labs and regulators globally may reference as formal pre-deployment evaluation norms develop.

Why it matters: Frontier AI labs now face a concrete precedent for government pre-deployment model access that regulators in other jurisdictions are likely to follow.

AI Stocks

(TSM) TSMC Q2 earnings drop today with +36% revenue beat, AI demand intact

TSMC reported Q2 2026 earnings on July 16, posting revenue at the top of its $39 to $40.2 billion guidance range and beating year-over-year consensus by roughly 36%, confirming that AI chip orders have not decelerated. The results arrived directly after a sector-wide semiconductor selloff that erased an estimated $1.3 to $1.4 trillion in market capitalization over a matter of days, making management's forward commentary on CoWoS advanced packaging capacity, the 2nm process ramp, and full-year capital expenditure guidance the most closely watched signal in the AI infrastructure investment cycle. TSMC has now beaten consensus estimates in each of its last eight quarters, with an average positive surprise of 8.34%.

Why it matters: TSMC's beat and forward guidance will directly reset AI infrastructure investment expectations for chipmakers, hyperscalers, and their investors.

(AMZN, GOOGL, META, MSFT, ORCL) Hyperscaler free cash flow turns negative as $1.8T AI capex overwhelms earnings

Amazon, Alphabet, Meta, Microsoft, and Oracle are projected to see their combined free cash flow fall from a 2024 peak of roughly $250 billion to approximately $100 billion by end-2026, driven by a projected $1.8 trillion in AI capital expenditure across 2026 and 2027. Oracle has already entered negative FCF territory at negative $23.69 billion for fiscal 2026, and Amazon's Q1 capex of $44.2 billion outpaced operating cash flow of $26 billion. Oracle CFO Hilary Maxson told analysts the company expects to raise around $40 billion in new debt and equity in fiscal 2027. BofA projects that the capital flowing out of hyperscaler balance sheets will land with chip suppliers, forecasting a record $430 billion in combined free cash flow for Nvidia, Micron, Broadcom, and AMAT over the next 12 months.

Why it matters: The cash inversion from hyperscalers to chip suppliers marks a structural shift that investors, AI builders, and enterprise buyers should treat as a long-term supply-chain signal.