AI-Proof - Weekly AI Pulse
A concise summary of the week’s most important AI developments
Executive Summary
Money, infrastructure and governance now matter as much as model performance. Nvidia is working with six financial groups to mobilise more than $500 billion for data centres, Databricks has raised $5 billion at a $190 billion valuation, and IBM is building an OpenAI practice around thousands of consultants. Governments are also deciding how advanced models should be tested and who should pay for the power, water and grid upgrades behind them.
The tools are becoming more practical for everyday work. Claude can now carry browser tasks across sessions, Grok Bot can run routines on its own cloud computer, and faster models are cutting waiting times sharply. For most businesses, the sensible move remains a narrow trial with limited permissions, clear measures and a person approving anything consequential.
What to Try This Week
Test Claude on one browser-based admin task
If you have a paid Claude plan, install Claude in Chrome and choose one trusted portal, such as invoicing, bookings or supplier management. Ask it to collect a defined set of records and place them in a spreadsheet. Keep automatic approval off, review every action and compare the result with doing the task manually. Measure time saved, errors and anything that required intervention.
Turn one repeat task into a Gemini Gem
If your team uses Gemini, create a Gem for one repeatable task: briefing notes, first-draft emails or document review. Give it a strong example, the required format, the source material it may use and the checks it must make. Add a clear stop point, such as “show me the draft and wait for approval”. Run it three times and note where the instructions fail.
Use Grok Bot’s free 1 week trial
If your account offers the Grok Bot free trial, give it repetitive workflows that normally takes at least an hour a week, such as checking a portal, compiling a report or updating a tracker. Use a separate low-permission account and require approval before any external action. Record setup time, weekly time saved and mistakes. If you would continue through Cursor Ultra, test whether the savings justify its $200 monthly price.
Geopolitics, Governance and Big Moves
White House may add open models to classified pre-release checks
The White House is expected to expand its voluntary pre-release testing framework once open-weight models match the strongest closed systems. They are currently excluded. The testing criteria are classified and participation is voluntary, leaving questions about triggers, enforcement and whether open developers could comply with a 30-day review. This is a reported plan, not an enacted rule.
Sources: Wired reporting, Axios framework analysis, AP background
Databricks raises $5 billion at a $190 billion valuation
Databricks has closed a $5 billion round led by Coatue, with backing from Blackstone, MGX, T. Rowe Price and Sixth Street. Its valuation has risen from $134 billion in February to $190 billion. The company says annualised revenue has passed $7 billion and is growing more than 80% year on year. The money will fund enterprise AI products, acquisitions and hiring while allowing Databricks to delay an IPO.
Sources: Forbes reporting, CNBC reporting
IBM builds an OpenAI practice around thousands of consultants
IBM has joined the top tier of OpenAI’s partner network and plans to certify thousands of consultants and engineers. The firms will combine OpenAI models, Codex and ChatGPT Work with IBM’s industry and delivery teams, focusing on business processes, software modernisation and cyber security. The announcement is mainly a services distribution deal, not an exclusive technology partnership.
Sources: Partnership reporting, OpenAI Partner Network
Estimates put OpenAI’s revenue run rate above $40 billion
Third-party estimates put OpenAI’s annualised revenue above $40 billion, up from $25 billion in February. OpenAI has not published audited results, and a run rate extrapolates one recent month rather than measuring full-year sales. The company has confirmed that business products now account for more than 40% of revenue. Treat the headline figure as an estimate, not a reported result.
Sources: Axios estimates, Reuters reporting on the February figure, OpenAI enterprise disclosure
Nvidia recruits Wall Street for a $500 billion infrastructure push
Nvidia has signed preliminary agreements with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to create independent financing platforms for AI data centres. The aim is to mobilise more than $500 billion of third-party capital, largely for Nvidia-based infrastructure. Nvidia is not putting up $500 billion itself. The plan could widen access to funding, but it raises concerns about debt and suppliers helping to finance their own demand.
Sources: Nvidia announcement, Axios analysis
OpenAI loses two senior commercial leaders in one week
Longtime executive Brad Lightcap is leaving to start something new, while revenue chief Denise Dresser will be replaced by Dali Rajic, formerly president and chief operating officer of Wiz. The changes follow Fidji Simo’s July departure for health reasons. Greg Brockman is taking more operating responsibility as OpenAI pushes harder into enterprise sales and prepares for a possible flotation.
Sources: Denise Dresser reporting, Brad Lightcap reporting
OpenAI courts US states as data centre resistance grows
OpenAI is urging states to align around a common AI safety framework while expanding Stargate data centres, including major Texas sites. Texas is also tightening infrastructure rules: Governor Greg Abbott wants operators to fund grid upgrades, reduce water use and disclose demand. OpenAI’s state engagement now concerns power, water and local economics as much as model regulation.
Sources: OpenAI state policy proposal, Texas governor’s directive, OpenAI infrastructure plans
Tools and Releases
Anthropic adds invisible watermarks to new Claude models
Claude models launched in the EU on or after 2 August 2026 now embed an imperceptible, machine-readable watermark in generated text. Anthropic says the marking applies worldwide across Claude, Claude Code, Cowork, its API and supported cloud platforms. Supported files, including SVG, PNG and JPG, receive signed C2PA provenance metadata.
Claude turns its Chrome side panel into a Cowork session
Claude’s Chrome side panel now carries the same Cowork conversation, skills and connectors across browser, desktop and mobile. It can handle multi-step work inside logged-in websites, but prompt injection remains the main risk. Anthropic advises using trusted sites, restrictive allowlists and manual approval for consequential actions. Availability is rolling out across paid plans, with Enterprise access controlled by administrators.
Sources: Anthropic safety guidance, Anthropic setup guide
OpenAI slows Astra after cyber tests trigger its highest risk threshold
OpenAI has slowed work on Astra, an unreleased model whose internal tests could not rule out critical cyber capability. The company is expanding evaluations and pausing internal use that cannot meet tighter controls. Calling Astra GPT-6 is unsupported, and no public test confirms the capability. OpenAI has applied its strictest cyber safeguards before release.
Source: Axios reporting
Grok 4.6 closes the gap on the leading models
SpaceXAI has released Grok 4.6 for coding, multi-step work and knowledge tasks. Headline API pricing is $2 per million input tokens and $6 per million output tokens, although rates rise for very long prompts. Independent aggregate testing places it close to GPT-5.6 Sol and Claude Fable 5, but rankings vary by test setup. It is available through SpaceXAI’s API and developer tools.
Sources: SpaceXAI announcement, independent comparison reported by Axios
Grok Bot puts always-on workers on their own cloud computers
Grok Bot lets users create AI workers that sign into web tools, follow recorded routines and hand work between bots while running in the cloud. The early beta is included with SuperGrok Heavy and Cursor Ultra, which costs $200 a month, with a team tier also available. Businesses should treat credentials, permissions, audit logs and approval points as launch requirements.
Source: Grok Bot product page
OpenAI previews a 14-times-faster GPT-5.6 Sol tier
OpenAI is previewing Ultrafast, an API tier that runs GPT-5.6 Sol on Cerebras hardware at up to 750 output tokens per second. OpenAI says this can be up to 14 times faster than standard processing. It changes response time, not the model’s intelligence, and access, capacity and pricing will determine whether it is practical beyond work where every second counts.
Source: OpenAI announcement
Microsoft starts merging its two Copilot apps
Microsoft has begun combining its consumer Copilot and Microsoft 365 Copilot apps, while keeping personal and work data separated by account. The rollout starts on mobile and web before reaching Windows and Mac. Consumer Deep Research and Podcasts close on 18 August. Microsoft 365 Premium users retain research through Researcher, but existing podcasts become inaccessible and Microsoft offers no export.
Sources: Fortune reporting, Microsoft Deep Research notice, Microsoft Podcasts notice
Google’s AMIE moves from text chat to live video consultations
Google researchers have tested a video version of AMIE that listens, speaks and interprets visual cues during simulated primary-care consultations. In a study involving 30 doctors, 15 patient actors and 100 scenarios, evaluators rated it on par with or better than doctors on several clinical measures. This was a controlled simulation, not evidence that the system is ready to diagnose or treat real patients.
Source: AMIE research paper
DeepSeek takes V4 Pro out of preview
DeepSeek has rolled V4 Pro’s production release across its app, website and API, with the existing API name now serving the latest version. It adds stronger performance on tool-using tasks, native Responses API support and low, high and max reasoning settings. DeepSeek’s published benchmark gains are company-run. Updated weights for this production build were not available when checked.
Source: DeepSeek changelog
Alibaba prepares open-weight Qwen 3.8 models
Alibaba has previewed Qwen3.8 Max and says it will release open weights for both the flagship model and a smaller 27B version. At the time of writing, the weights, model card and independent benchmarks were not yet public. Treat claims about scale or performance as provisional until researchers can download the models, inspect the licence and run like-for-like tests.
Source: Qwen announcement
ChatGPT can remember selected activity on your Mac
Computer History is an opt-in macOS feature that lets ChatGPT and Codex refer back to selected activity across apps and websites. It records interaction events, not screenshots, screen recordings or audio, and users can choose sources, pause collection and delete timeline items. It is off by default for Pro, Business and Enterprise users and is not yet available in the UK, EEA or Switzerland.
Source: OpenAI release notes
Google launches Gemini 3.7 Flash for coding and agents
Gemini 3.7 Flash is Google’s new high-volume model for coding, multi-step tasks and knowledge work. Introductory API pricing is $0.75 per million input tokens and $3.75 per million output tokens until 31 December, then doubles. It is available through Google’s developer and enterprise platforms, and powers Gemini Spark for eligible Pro and Ultra subscribers.
Source: Google announcement
Z.ai launches GLM-5.3, but holds back the weights
Z.ai says GLM-5.3 improves complex coding and long-running tasks through post-training, using the same base model as GLM-5.2. It also reports unexpectedly strong vulnerability discovery and exploitation results, so the company is holding back the weights for two weeks while it completes safety work. The benchmark claims are Z.ai’s own and need independent testing.
Source: Z.ai announcement
Quick Hits
OpenAI Tests Ads in ChatGPT – To sustain free access, OpenAI experiments with clearly labeled ads within ChatGPT conversations, maintaining response impartiality and user controls.
Nvidia release nemotron 3.5 Lightening a lightweight open 30B-parameter (MoE-style) model, optimised for agentic tasks.
Google’s Gemini app reportedly surpassed 1 billion monthly active users.






