VERIFIED AI SIGNALS
AI Industry News
Source-verified updates on AI models, APIs, research, security, and regulation.
Z.ai says the GLM-5.3 API is now available for software development, defensive cybersecurity, and long-horizon agentic workflows. The company says pricing matches GLM-5.2 and access is offered through its official API and partner gateways.
A post attributed to a Codex engineer says GPT-5.6 Sol can now use a one-million-token context in Codex sessions authenticated with ChatGPT accounts, not only API keys. The post also gives configuration and command-line examples.
Google has introduced Gemini 3.8 Flash for long-horizon coding and autonomous agents, alongside Gemini 3.8 Flash Cyber for vulnerability detection and automated remediation.
OpenAI’s GPT-5.6 Builder’s Guide explains how model routing, reasoning controls, and new Responses API capabilities can reduce the cost and latency of agent workflows. In the company’s BrowseComp example, GPT-5.6 Luna achieved 84.04% at very high reasoning for a launch cost of $1.33, close to GPT-5.5’s 84.36% result at $33.27.
OpenAI is previewing Ultrafast, a limited-access API tier for GPT-5.6 Sol. Powered by Cerebras, it is advertised at up to 14× Standard processing speed and up to 750 output tokens per second.
MiniMax announced a joint livestream with fal focused on building with MiniMax H3, including multimodal references, native audio, editing, LoRA training, model chaining, and production workflows. Supporting materials describe H3 as an open-weights multimodal video model available through fal.ai’s hosted API.
Google has introduced Gemini 3.7 Flash with reported gains in software engineering, knowledge-intensive work, web development, and agent workflows. The model also debuts at an introductory price equal to half the original Gemini 3.6 Flash cost per million tokens.
OpenAI is previewing an API service tier for GPT-5.6 Sol that is powered by Cerebras and claims output speeds of up to 750 tokens per second.
xAI launched Grok 4.6 on Aug 12, 2026, scoring 61 on the Artificial Analysis Intelligence Index and matching GPT-5.6 Sol Max. API pricing remains $2/$6 per million input/output tokens, context expands to 500K, and a new xhigh reasoning tier is available. The model is live across xAI API, Grok Build, Cursor, OpenRouter, Vercel, and Cloudflare.
Databricks says Moonshot AI’s open-weight Kimi K3 model is now available through Unity AI Gateway. Organizations can access it within their existing Lakehouse environment while applying centralized controls for permissions, monitoring, and governance.
DeepSeek has launched the stable V4-Flash API in public beta, reporting substantial gains on agent benchmarks, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon Verified.
DeepSeek V4 Flash appears in several model directories and comparison pages, with reported advantages in context length, speed, and cost. Available evidence does not confirm all ColaOS launch details, and a third-party benchmark comparison contradicts the claim that it is more intelligent than GLM-5.2.
Several secondary sources report that Alibaba’s Qwen team has released Qwen3.8-Max and plans to publish its model weights the following week. The model is described as a multimodal flagship with roughly 2.4 trillion total parameters and 95 billion activated parameters.
MiniMax H3 has been integrated into vLLM-Omni alongside the release of its open weights. The serving framework exposes OpenAI-compatible asynchronous and synchronous video-generation endpoints.
DeepSeek has opened a public beta for the official V4-Flash API. The release keeps the existing model name, improves performance across several agent benchmarks, and adds native Responses API and Codex support.
Alibaba’s Qwen has introduced Qwen3.8-Max, describing it as the most capable model in the Qwen family so far. The model is reported to have 2.4 trillion total parameters, 95 billion activated parameters, and a 1-million-token context window.
OpenAI says GPT-5.6 Sol helped optimize the production infrastructure and inference stack used to run it. The work spans GPU kernels, routing, scheduling, caching, and model implementation; a secondary report attributes roughly 20% lower serving costs and more than 15% better token-generation efficiency to the changes.
DigitalOcean has added Moonshot AI’s Kimi K3 to its Inference Engine, allowing developers to access the model through managed Serverless Inference without operating their own serving infrastructure.
DeepSeek has released the V4-Flash-0731 API in public beta. The update focuses on agentic and coding performance, adds native OpenAI Responses API support, and enables a more direct Codex integration.
One user says they consumed roughly 2.8 billion tokens with Opus 5 in a day and found it better than the revised Fable 5 across most capabilities. The same account gives GPT-5.6 Sol an edge on complex backend logic and long-running tasks. Available evidence supports Opus 5's efficiency, but sources disagree about its overall ranking against the two competing models.