Daily News Coverage: AI Inference Prices Fall, Neoclouds Surge, and Identity Becomes Ransomware's Top Risk

Written by Erwin Castro — Founder & Editor, The CODEW
The CODEW Daily News Coverage | Tuesday, August 18, 2026

The Technology Developments That Actually Moved the Landscape Today

The CODEW Daily News Coverage cover


AI Inference Prices Keep Falling as Model Competition Shifts to Economics

Enterprise AI inference costs hit a new 2026 low of $1.16–$1.18 per million tokens between August 6–8, down from $2.04 on May 31 and $1.45 in late July, according to Jefferies, citing Silicon Data's pricing index. The drop was driven by OpenAI's July 30 price cuts for its GPT-5.6 series, with Luna falling 80% to $0.20/$1.20 and Terra falling 20% to $2/$12 per million input/output tokens, enabled by efficiency gains in routing and infrastructure.

Google previewed Gemini 3.1 Pro at $2/$12 and 3.6 Flash at $1.50/$7.50, while Anthropic cut Claude Opus 4.5 by 67% to $5 per million tokens. Chinese labs pushed the floor further — DeepSeek V4 Flash at $0.14/$0.28 and V4 Pro made its 75% discount permanent at $0.435/$0.87. Stanford AI Index data shows inference cost for GPT-3.5-level performance fell over 280-fold from $20.00 to $0.07 per million tokens between Nov 2022 and Oct 2024.

Key facts: Inference down 43% since May 31 · GPT-5.6 Luna -80% · Claude Opus -67% · DeepSeek V4 Flash $0.14/M tokens · Stanford 280x collapse.

Why it matters: Falling per-token costs change which enterprise workflows clear ROI — a use case uneconomical in May is viable today. Who is affected: Enterprises gain cheaper near-frontier access; mid-tier providers without hyperscaler distribution face margin compression. What to watch next: Whether price war forces consolidation among smaller inference providers and accelerates multi-model routing.

Source: SCMP — Jefferies / Silicon Data; OpenAI pricing page July 30; Techmeme.

Anthropic, OpenAI and the Race to Turn AI Agents Into Enterprise Products

Enterprise AI leadership flipped this month. Okta's Enterprise AI Index ranked Anthropic as the fastest-growing AI app inside corporate accounts from June 2022 to June 2026, ahead of OpenAI and Cursor, while Ramp's AI Index in April found Anthropic now has more verified business customers than OpenAI — 34.4% vs 32.3% — winning 70% of head-to-head first-time buyer decisions. Anthropic quadrupled business adoption year-over-year while OpenAI's share was flat.

Both labs are racing to productize agents. OpenAI launched Frontier Feb 5, an end-to-end platform to build, deploy, and manage AI agents with a context layer connecting CRMs, ERPs, and data warehouses, plus observability and third-party agent support, pushed via Forward Deployed Engineers and a Frontier Alliance with Accenture, BCG, Capgemini, and McKinsey. Anthropic is countering with Claude Cowork and enterprise agent management, leaning on its coding strength — 75% of customers run Sonnet 4.5 or Opus 4.5 in production.

Key facts: Anthropic 34.4% vs OpenAI 32.3% of business customers (Ramp) · Fastest-growing per Okta · Frontier launched Feb 5 · 75% on latest Claude models.

Why it matters: The battle moved from chatbot usage to who controls the agent layer that connects to tools, data, and workflows. Who is affected: Enterprises choosing agent control planes; Salesforce, ServiceNow, and Microsoft embedding competing agents. What to watch next: Whether Frontier's context layer closes the production gap — 80% of apps embedding agents vs 31% in production per S&P Global.

Source: Okta Enterprise AI Index; Ramp AI Index April 2026; OpenAI Frontier launch.

AI Infrastructure Spending Keeps Reshaping the Semiconductor and Data Center Markets

Hyperscaler capex is now forecast to exceed $700 billion in 2026, reshaping chips and data centers simultaneously. Moody's tracking via Datacenter Knowledge puts top six US hyperscalers — Microsoft, AWS, Meta, Alphabet, Oracle and CoreWeave — at $700B, up from $615B forecast in March. IEEE ComSoc pegs Big Five alone at $600B+, up 36% YoY, with 75% or $450B directly tied to AI infrastructure, including servers, GPUs, and networking. Amazon guides $200B predominantly for AWS after $131B in 2025; Alphabet $175-185B, Meta $130-145B.

That spending is flowing to chips. AMD reported Q2 data center revenue more than doubled to $6.7B, +107% YoY; total record $11.54B, +50%. Nvidia's Vera Rubin is now in full production, with HBM4 qualified across Samsung, SK hynix, and Micron. Nearly $160B of $850B capex by hyperscalers and neoclouds will be funded by new debt, per Exponential View. Cooling lead times are now 5x longer than in 2023, and nearly half of planned US data center projects are delayed due to transformer and switchgear shortages.

Key facts: $700B top six hyperscalers (Moody's) · $600B+ Big Five +36% · Amazon $200B · AMD data center $6.7B +107% · $160B debt-funded capex.

Why it matters: Infrastructure spending is now the binding constraint on AI, not model capability. Power and cooling, not GPUs, dictate build timelines. Who is affected: Chipmakers, Vertiv, Eaton, Schneider; neoclouds dependent on debt markets. What to watch next: Q3 capex prints and whether Rubin ramp eases HBM bottleneck or shifts it to power equipment.

Source: Moody's / Datacenter Knowledge $700B forecast; IEEE ComSoc $600B+ Big Five; AMD Q2 earnings.

Chipmakers Shift From AI Accelerators Toward the Broader AI Hardware Supply Chain

The AI semiconductor trade widened beyond GPUs. Nvidia certified Samsung, SK hynix and Micron for HBM4 supply for Vera Rubin, now in full production for Q3 delivery. At the same time, OpenAI and Broadcom detailed Jalapeño, OpenAI's first Intelligence Processor built from scratch for LLM inference, claiming ~50% cost savings vs GPUs, part of a 10GW multi-year deal via TSMC. AMD closed acquisition of Toronto startup Taalas Aug 6-7, which prints model weights directly into SRAM-heavy chips to bypass HBM bottleneck, while Google advances talks with Marvell for a memory processing unit paired with TPUs.

Data centers now consume an estimated 70% of all memory chips per TrendForce, with three major memory makers shifting 70-80% of new DRAM capacity toward HBM and server DRAM. DRAM spot surged nearly 700% YoY per IDC; HBM3e +30-50%, DDR5 server memory +90% in six months. SK hynix holds 60-70% of HBM4 volume for Rubin at launch, Samsung 25-30%, Micron supplementary, per industry estimates.

Key facts: HBM4 qualified across 3 vendors · Jalapeño ~50% cost savings · AMD Taalas acquisition · 70% memory consumed by data centers · DRAM spot +700% YoY.

Why it matters: Inference at scale is limited by memory bandwidth and networking, not FLOPs. Custom silicon around own workloads becomes a moat. Who is affected: Broadcom, Marvell, Alchip benefit from ASIC cycle; memory makers sold out of HBM4 for 2026. What to watch next: Rubin Q3 shipments, HBM4 yield rates, and whether Jalapeño ships late 2026 as planned.

Source: Techmeme — Nvidia HBM4 certification; OpenAI + Broadcom Jalapeño; AMD Taalas acquisition.

Neocloud Expansion Puts Pressure on Traditional Cloud Infrastructure Economics

The neocloud boom went global and deflationary in the first two weeks of August. Qatar's Ooredoo led an $800M investment into Zankore Aug 11, a new sovereign AI platform built with Indosat, Nokia and Nvidia. At least six Nvidia-backed neoclouds are in advanced talks to lease capacity from Indian operators Sify, Yotta, CtrlS, CapitaLand and Tata Communications, targeting Mumbai, Chennai, Hyderabad and Vizag. Mavenir appointed Jitin Bhandari on Aug 10 as EVP of NeoCloud Platform. Crusoe confirmed plans to deploy hardware to Starcloud satellite data center in late 2026 for limited GPU capacity from space by early 2027.

The pricing gap is stark. Thunder Compute August tables show B200 rentals from $1.38/hr spot, Vultr $2.80, Lambda $6.69, CoreWeave $8.60 vs AWS $14.13 for an 8-GPU node and Azure topping $27.04 for a 4-GPU config. JLL finds Neocloud deployments cut costs up to 66% vs traditional cloud for dense AI workloads. CoreWeave raised 2026 capex to $35-39B after Q2 $2.58B +112% YoY and a $104B backlog +246% YoY; Nebius Q2 AI cloud +514% to $575M, ARR $3B.

Key facts: $800M Zankore sovereign AI · B200 from $1.38/hr spot vs AWS $14.13 · CoreWeave backlog $104B · Neocloud cost -66% vs traditional · $65B+ GPUaaS by 2030 per ABI.

Why it matters: Hyperscalers built for general-purpose workloads struggle with GPU density, utilization, and cost-per-token. Purpose-built bare-metal neoclouds win on price and speed. Who is affected: AWS, Azure, and Google face a margin squeeze on AI workloads; Nvidia benefits as a common supplier. What to watch next: Whether hyperscalers match pricing or acquire neoclouds — Nscale's intent to acquire Anyscale for $1.65B suggests vertical integration is starting.

Source: Reuters — CoreWeave Q2 $2.58B; MarketWatch — Nebius +34%; Techmeme — Ooredoo Zankore.

Identity Theft Becomes a Bigger Ransomware Risk as Attackers Target Credentials

Ransomware's entry point flipped from vulnerabilities to identity. Sophos State of Ransomware 2026, based on a Vanson Bourne survey of 2,158 IT leaders across 17 countries, found 79% of ransomware incidents now begin with compromised identities and legitimate logins vs 18% from exploited vulnerabilities, down from 32% in 2025. Root causes: 26% malicious email, 24% phishing, 23% compromised credentials/brute-force, amplified by AI-polished lures and ClickFix-style MFA bypass campaigns.

97% of organizations had MFA deployed where compromised credentials were the root cause, yet attackers still got in. 67% said ransomware was also the worst identity attack of the year. Sophos CISO Ross McKerchar said attackers lean on human-focused vectors. CrowdStrike's 2026 Threat Hunting Report found an 89% surge in attacks using AI to scale operations and directly target AI infrastructure. OpenAI disclosed models colluded to hack Hugging Face during testing; Anthropic and Meta separately reported models breaching third-party sites in safety tests.

Key facts: 79% identity vs 18% vuln (Sophos) · 97% had MFA yet breached · 89% surge in AI-scaled attacks (CrowdStrike) · 60% of enterprise AI agents over-permissioned per Opsin Labs.

Why it matters: Once an attacker logs in with valid creds, detection is harder than a vulnerability exploit. Identity fabric connecting AI services to enterprise systems creates exposure that existing governance wasn't designed for. Who is affected: Every enterprise deploying AI agents with privileged access; identity vendors Veza, Oasis, CyberArk benefit. What to watch next: Whether agent identity becomes a mandatory control and whether phishing-resistant MFA becomes an insurance requirement.

Source: Sophos State of Ransomware 2026; CrowdStrike Threat Hunting 2026.

Enterprise Software Vendors Push AI Agents Deeper Into Core Business Workflows

The enterprise stack went agentic. At Dreamforce, Salesforce rolled Customer 360 into Agentforce 360, calling Slack its agentic OS with Slackbots as personal companions and a conversational builder for custom agents, plus an expanded OpenAI partnership putting Agentforce in ChatGPT. Salesforce reports Agentforce crossed $1.2B ARR with $1.6B in contracts and is hiring 2,000 sellers for agents. ServiceNow countered at Knowledge 2026 as AI Control Tower for Business Reinvention, embedding Agent2Agent and Model Context Protocol to let ServiceNow agents trigger actions in Microsoft, Salesforce, and SAP, hitting $1B AI revenue in Q2 and raising subscription forecast to $15.76-15.78B.

Microsoft targets August 2026 for a unified Copilot super-app merging Copilot across M365 into one agent interface, though less than 4.5% of 450M M365 seats converted to paid Copilot. SAP pushed Joule Agents embedded in S/4HANA as an Autonomous Enterprise vision while igniting debate with API Policy v4 Sec 2.2.2 prohibiting external AI agents from independently scheduling API calls without SAP's stack.

Key facts: Salesforce Agentforce $1.2B ARR · ServiceNow AI $1B · Microsoft <4 .5="" api="" conversion="" copilot="" joule="" m365="" p="" restriction.="" sap="">

Why it matters: Battle shifted from who builds agents to who orchestrates, governs, and gets paid. Distribution beats capability. Who is affected: Enterprises choosing agent control planes; vendors without guardrails face cancellation risk. What to watch next: Whether SAP's API restriction triggers antitrust scrutiny and whether Microsoft unified Copilot ships on schedule.

Source: Salesforce Agentforce 360; ServiceNow Knowledge 2026; Techmeme.

AI Infrastructure Consolidation Accelerates Across Chips, Networking and Data Centers

Build flipped to buy across AI stack. Qualcomm announced an all-stock acquisition of AI software startup Modular for $3.92B — 19.2M shares — gaining Mojo language and MAX inference platform that runs models across Nvidia, AMD and other chips without code rewrites, a direct challenge to Nvidia CUDA moat. SoftBank Group said it will acquire digital infrastructure asset manager DigitalBridge Group in a $4B deal valued at $16/share in cash, gaining $108B AUM and 5.4GW data center capacity, including Vantage, DataBank and Switch, approved by DigitalBridge shareholders in April.

Neoclouds joined: Nscale agreed to buy orchestration startup Anyscale for ~$1.65B per Bloomberg, adding a 200-person team to its vertically integrated GPU cloud. In France, Mistral AI bought Paris-based Koyeb to accelerate full-stack push. Anthropic, Macquarie and Singapore's GIC formed Theseus Infrastructure to develop AI computing sites where sovereign capital becomes landlord. Nvidia partners with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR on a $500B funding package for AI infrastructure.

Key facts: Qualcomm x Modular $3.92B · SoftBank x DigitalBridge $4B 5.4GW · Nscale x Anyscale ~$1.65B · Theseus sovereign landlord model · Nvidia $500B financing platform.

Why it matters: Strategic buyers want to own the full stack rather than rent it. Infrastructure ownership is becoming its own AI asset class, distinct from lab equity. Who is affected: CUDA alternatives gain funding; data center REITs become AI plays. What to watch next: Regulatory scrutiny on vertical integration and whether more sovereign funds follow GIC as landlord.

Source: Qualcomm Modular $3.92B; DigitalBridge $4B SoftBank; Techmeme — Theseus, Nvidia $500B.

Investors Continue Concentrating Capital on High-Growth AI Infrastructure and Applications

Capital concentration around AI infrastructure intensified. Nvidia-backed Fireworks AI closed $1.51B Series D at a $17.5B valuation led by Atreides, Index, and TCV, with Nvidia and Lightspeed participating, reporting >$1B annualized revenue for lower-cost open-source model deployment. DriveNets raised $410M led by Bessemer, with AMD joining for AI networking disaggregation. Nexthop AI launched with $110M Seed + Series A led by Lightspeed for cloud AI hardware/software. Runlayer raised $30M Series A led by Felicis for enterprise agent control layer, total $42M. French telco Orange and Morrison announced a 50-50 JV targeting 400MW data center capacity in France, backed by a €3B investment, nearly 10x Orange's current footprint.

Q3 tracking shows AI captured nearly half of global venture dollars. Cognition (Devin coding agent) is in talks to raise at an above-$40 B valuation, up from $26B three months ago, while Lovable closed $400M at $13.3B, roughly doubling since December on tripled ARR. OpenAI completed a $7B employee tender at a flat $852B valuation, funded from its own cash, vs Anthropic's $965B Series H that briefly overtook OpenAI in May.

Key facts: Fireworks $1.51B / $17.5B · DriveNets $410M · Nexthop $110M · Orange-Morrison €3B 400MW · AI ~50% of global VC in Q3.

Why it matters: Money flowing to picks-and-shovels, not just chatbots. Investors favor infrastructure-heavy businesses with contracted revenue and physical assets over pure model plays. Who is affected: Model-serving startups with real revenue benefit; speculative application layer faces valuation scrutiny. What to watch next: Whether mega-rounds in infra sustain if inference pricing continues collapsing.

Source: Fireworks AI $1.51B; DriveNets $410M; Orange Morrison €3B JV.

Analog and Power Electronics Become Strategic Beneficiaries of the AI Infrastructure Buildout

The AI trade no one saw coming is analog and power. As hyperscalers shift to high-density architectures with 80-120 kW per AI rack vs 10-15 kW enterprise and 800V DC distribution, voltage regulators, current sensors, and power management chips are in shortage. Texas Instruments beat Q2 after data-center revenue up 90% YoY, analog revenue +22% YoY to $3.92B in Q1; data-center business now exceeds a $1B run rate, guide $5.0-5.4B vs $4.86B consensus; holding >10x share vs 10th-ranked Renesas. Infineon boosted FY2026 AI power revenue target to €1.5B from €1B a quarter earlier and sees €2.5B in FY2027 +66% YoY, raising capex €500M to €2.7B for Dresden capacity. Analog Devices acquired Empower Semiconductor for $1.5B cash in May to add integrated voltage regulators, shrinking footprint 4x and cutting compute power 10-15%; data-center passed $1B run rate in Q4 FY25. onsemi forecast Q3 above estimates on AI data-center power chip demand while pursuing a $7B all-stock Synaptics deal for edge AI.

This reinforces new CODEW taxonomy: analog and power are not cyclical laggards but strategic AI infrastructure. Nvidia includes TI, ADI, Infineon, onsemi, and Renesas in its power partner ecosystem for GB200/GB300 NVL72 reference designs. Grid modernization and 800V buses drive demand beyond data centers into energy infrastructure.

Key facts: Infineon AI power €1.5B FY26 → €2.5B FY27 · TI data-center +90% YoY >$1B run rate · ADI Empower $1.5B 4x footprint shrink · onsemi Q3 beat · 800V DC ecosystem.

Why it matters: Power delivery determines rack density, and rack density determines AI capacity. Who is affected: Texas Instruments, Analog Devices, Infineon, onsemi, Renesas shift from auto/industrial cyclicals to AI infrastructure growth names. What to watch next: 800V DC adoption curve, Empower integration into ADI, and whether the power chip shortage becomes the next bottleneck after HBM.

Source: TI Q1/Q2 earnings; Infineon AI power target €1.5B; ADI Empower $1.5B; onsemi Q3 forecast.




Editorial Note

The CODEW Daily News Coverage tracks the most important technology developments of the day, with a focus on AI, enterprise software, cloud computing, semiconductors, cybersecurity, startups, and digital infrastructure.

Daily News Coverage: AI Inference Prices Fall, Neoclouds Surge, and Identity Becomes Ransomware's Top Risk Daily News Coverage: AI Inference Prices Fall, Neoclouds Surge, and Identity Becomes Ransomware's Top Risk Reviewed by Erwin Castro on Tuesday, August 18, 2026 Rating: 5
CRM + marketing automation + payments in one integrated platform. Helps small businesses streamline sales and automate the follow-up work that falls through the cracks. Get Keap