Virexa
HomeAIProgrammingCloudSecurityOpen SourceGamesMobile GamesDeveloper Hub
Sign InSign Up
Virexa
Sign InSign Up
AIProgrammingCloudSecurityOpen SourceGamesMobile GamesDeveloper Hub
Virexa

Modern AI news aggregation and newsletter platform covering technology, business, AI, games and world news.

Categories

  • AI
  • Programming
  • Cloud
  • Security
  • Open Source
  • Developer Hub

Company

  • About
  • Contact
  • Advertise

Resources

  • RSS Feed
  • API
  • Privacy Policy
  • Terms of Service

© 2026 Virexa. All rights reserved.

Virexa
HomeAIProgrammingCloudSecurityOpen SourceGamesMobile GamesDeveloper Hub
Sign InSign Up
Virexa
Sign InSign Up
AIProgrammingCloudSecurityOpen SourceGamesMobile GamesDeveloper Hub
Home›News Explorer

Explore

News Explorer

Browse every article collected by VIREXA. Newest articles appear first.

Filters

77 results • Page 2 of 7

AI coding agents are blowing through budgets — Replit, Kilo Code, and Symbotic explain how they're managing it
ProgrammingTutorial

AI coding agents are blowing through budgets — Replit, Kilo Code, and Symbotic explain how they're managing it

At Kilo Code, engineers are reading or writing code themselves only about 1% of the time now, according to co-founder Emilie Schario — the rest is agents. That shift is forcing new questions onto dev teams: which systems are safe to hand over, who cleans up when models goof up, how to support multi-model architectures, and whether skyrocketing token bills mean real progress or just burned IT budget. As far as tech leads from Replit, Kilo Code, and Symbotic are concerned, it’s a natural — and welcome — evolution as agentic AI becomes embedded into more and more enterprise workflows. “Unless something's really broken or debugging, 99% of the time engineers are not reading or writing code anymore,” Emilie Schario, co-founder of Kilo Code, said at VB Transform 2026 . AI good at greenfield, not so great at brownfield For Jared Go, distinguished engineer for AI and cloud at warehouse automation company Symbotic, the current moment is about directing the focus of AI. "These are my criteria," he said. "Let's look at it from the lens of security, elegance, clean, concise code, water tightness." That way, AI does most of the heavy lifting, and human code review isn't as critical. Human involvement becomes necessary further down the line, Go noted, because agents don't make strong product decisions. “Greenfield [building brand new codebases] is so easy for agents. Brownfield [writing, updating, or maintaining existing code] we all know is where the actual challenge lies.” Replit takes a bit of a different tack: While the company has "gone very agentic," they've been more conservative with AI coding, explained Amol Jain, head of product engineering. An agent reviews each pull request (PR) and assigns it a risk score; low-risk PRs are self-merged by their author, while others go to human reviewers who read the code and give feedback. “The idea was human on the loop, not human in the loop,” Jain said. Replit’s internal tool is essentially self-driving for software engineers; devs give a task to agents, which do end to end planning, implementation, and testing. “It's a fleet of agents that run in their own cloud virtual machines (VMs) with access controls behind token proxies so they're secure,” Jain said. He shared one example where an engineer couldn’t repro or solve a “very gnarly bug” deep in its systems. It was sent to an AI manager agent, which told it to go to sleep. The manager agent then spun up a bunch of underlying agents that found the issue; it subsequently spun up a bunch more agents that found the fix. Six hours later, AI had a PR ready for the bug that had puzzled human engineers. Multi-model is the future AI providers are also evolving beyond the lock-in model, as customers increasingly demand multi-model choice. Kilo Code, for its part, supports 500-plus models in its gateway. "Your software that you're using to do agentic engineering should be decoupled from the model that you're using to do it," Schario said. For instance, Schario said companies often use expensive frontier-tier models to architect a project, then switch to a less expensive open-weight model for the rest of the work. It’s also important to respect model provider limitations, such as when they need to work in closed or isolated environments or providers in their specific regions. “It's factoring in what's important to you, what limitations you've set, what data retention policies you've established, what keys you've brought in, what commits you might have … into that routing decision,” Schario said. Replit, similarly, tends to have a better sense of the cost versus capability spectrum than its customers, Jain contended. “We are essentially making the decisions on users' behalf of what model to use when, in what capacity, to minimize cost and maximize capability.” To tokenmaxx or not to tokenmaxx Of course, an important consideration as AI adoption increases is runaway costs, which has led to some enterprises tracking and capping AI use through tokenmaxxing. Concerns come from both sides, Schario said: internally and from customers. From the latter, she's hearing, "I accidentally spent my whole AI budget for the year … so what do I do now?" In response, Schario said Kilo Code points customers to the same workflow: use expensive models for planning, then open-weight models for affordability. Further, sharing skills, strong guidance, and Model Context Protocol (MCP) will empower models. “Realizing where you can really uplevel your team to help them get the most out of the models they're using is going to make a big difference,” Schario said. Internally, meanwhile, Schario noted one particular engineer that has a "heavy foot" and is constantly at the top of the usage board. "I regularly have to nudge, 'What are you doing there?'" she said. It's easy to look at a $600 bill for daily work and react, "Wow, that's so much," but looking at the amount of work completed can sometimes justify the cost. “Cost per pull request is the metric that I'm paying attention to right now,” Schario said. “It feels like the closest proximity for how I can measure value.” Ultimately, AI changes how enterprises are thinking about ROI because spend is not the problem. “The spend with no return on that spend is the problem.” Symbotic, for its part, has set per-month cost tiers for its employees. The company built a tool that gives managers visibility into PRs and usage trends. They can then move users up or down a tier as they see fit, Go explained. “Having a cap and seeing how many people went up in cap this month makes a big difference when you're trying to corral these costs and make things efficient,” Go said. When Cursor — which Symbotic uses heavily — ended a legacy discount that had grandfathered the company into a flat per-request rate even for frontier models, and moved everyone to full pricing, it forced a company-wide reckoning on efficiency, Go said. "People were saying, 'You should try this model … This works better for this C# code, this whatever,'" he said. But the cost problem is increasingly moving out of IT; Replit, for one, broadened agents beyond engineering, and eventually found that a user on the support side had "blown through an insane amount of money," Jain said. When they looked under the hood, they figured out it was because they were running an automation on GPT 5.5 Pro Max. “At least till that point, the ROI was rather clear,” Jain said. “We could see engineering productivity 3X, so no one had questioned it yet.” Visibility that isn’t “anti-productive,” model routing, and sensible defaults are critical, he emphasized. “Most tasks do not need the frontier.”

← Previous1234…7Next →
🔥

Developer Pulse

What developers are discussing today

  • GPT-5.5 API↗9.4K
  • Next.js 16↗6.2K
  • Claude Code↗5.8K
  • Kubernetes→3.4K
  • Rust↗2.7K
VentureBeat·August 4, 2026·6 min read
Commerce AI has a measurement problem no one is talking about
AIResearch

Commerce AI has a measurement problem no one is talking about

Presented by Rezolve Ai Most brands know something is shifting in how consumers find and choose products. What most don't know is how much of that shift has already taken place, where it's happening, or whether they're on the right side of it. That uncertainty is the problem. And the analytics stack most brands rely on isn't built to resolve it. The decision layer has moved In 2014, 82% of digital commerce started on a brand's website. By 2024 that had fallen to 38%, according to Salesforce research . The journey that used to begin at a brand's front door now begins somewhere else. Increasingly, it begins with a question asked of an AI platform and ends with an answer that shapes the purchase decision before any brand-owned touchpoint is engaged. Consumers are asking AI where to shop, what to buy, and which product is right for them. Bain research shows that four in five consumers rely on zero-click results at least 40% of the time . That means the shortlist a consumer receives from an AI answer engine is, in many cases, the only shortlist they consult. Adobe Analytics recorded over 800% year-over-year growth in AI-driven traffic to retail sites , a signal of how rapidly AI platforms are inserting themselves between brands and their customers. This is a structural shift, not a trend. And it has created a category of commercial loss that most analytics tools are architecturally incapable of detecting. What you can't see is costing you The gap is this: a brand can have strong onsite conversion metrics and still be losing significant ground in the market, because the customers who never arrived aren't captured in any dashboard. There's no "AI excluded you" event in a session log. There's no abandoned cart entry for a shopper who was told by an AI assistant that a competitor was the better fit. This is different from the SEO problem brands have managed for two decades. With traditional search, absence had a visible signal. You could see your ranking, audit the gap, and act on it. With AI answer engines, absence is invisible by default. The surface doesn't show you what it didn't show the consumer. Sixty percent of searches now end without a click, according to Semrush's 2025 zero-click study. For AI-mediated discovery, that number is structurally higher. The answer is the destination. If a brand isn't in the answer, it isn't in the consideration set, and its analytics will never surface that fact. The metric that isn't being measured The commerce industry has developed sophisticated instrumentation for the journey from landing page to purchase. It has essentially no instrumentation for the journey from consumer intent to brand discovery, the layer where AI is now operating. Brands that want to understand their actual competitive position in an AI-mediated market need to ask a different set of questions: How does my brand appear when consumers ask AI for recommendations in my category? What language does AI use to describe my products? Where am I present, where am I absent, and where am I being described in ways that don't reflect my positioning? These aren't marketing questions. They're infrastructure questions. And answering them requires a different kind of audit than anything in the current commerce or marketing toolkit. Rezolve Ai commissioned research across 1,500 US consumers in January 2025 that found the majority of shoppers who use AI for product research make purchase decisions directly from those AI-generated recommendations, without returning to a search engine or brand site to verify. The implication for brands is significant: by the time a consumer reaches a brand's owned properties, the decision may already have been made, or unmade, somewhere else. What comes next The brands that will maintain commercial relevance as AI mediates more of the discovery layer are those that develop visibility into it, not just presence on their own platforms. That means treating AI discoverability as a measurable discipline, not an assumption, and building the infrastructure to understand, track, and influence how AI systems represent them to consumers. The tools to do that are emerging. The measurement frameworks are not yet standardized. But the brands that begin building that visibility now will have a structural advantage as the market continues to shift. AI answer engines are already forming preferences. Every day without visibility is a day those preferences solidify without you. Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com .

VentureBeat·August 4, 2026·4 min read
Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use
AICertification

Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use

Chinese e-commerce and cloud giant Alibaba's famed Qwen team of AI researchers last night unveiled Qwen3.8-Max , a new flagship 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM) that targets one of the most competitive corners of the frontier AI market: autonomous software engineering and long-horizon enterprise work. If the company's published benchmarks hold up under broader independent testing, Qwen3.8-Max doesn't merely compete with today's leading proprietary models — it surpasses several of them on some key benchmarks in agentic computing. Most notably, Qwen reports that Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark measuring how well agents can use a computer operating system and applications on it, ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0). It also posted the highest reported score on PaperBench , the benchmark from OpenAI measuring how well agents can reconstruct scientific research papers from experimental data, and leading or remaining highly competitive across software engineering, research reproduction, multimodal reasoning, and visual web development benchmarks. The release also signals a potentially significant strategic shift for Alibaba: the company says open weights for Qwen3.8-Max will be released next week, alongside Qwen3.8-27B. If that happens under a permissive license, it would represent the first time a Max-class Qwen model becomes available for self-hosted deployment—a move that could substantially reshape enterprise adoption. One important caveat remains, however: Alibaba has not yet disclosed the licensing terms, leaving open the possibility that the release could use a more restrictive custom license, as we saw recently with Chinese rival Moonshot's open Kimi K3 frontier model , rather than a broadly permissive one such as Apache 2.0. A different definition of 'frontier' Over the past year, the competitive landscape for foundation models has become increasingly specialized. OpenAI has largely focused its GPT series on general reasoning, multimodal interaction and enterprise productivity. Anthropic's Claude series has emphasized coding and dependable long-context reasoning. Google continues to push Gemini toward multimodal productivity and web-native workflows. Moonshot AI's Kimi K3 recently entered the conversation by pairing frontier-class performance with an open-weight release. Qwen3.8-Max attempts to combine many of these strengths into a single model aimed squarely at enterprise automation. Rather than emphasizing conversational intelligence, Alibaba is positioning the model as an autonomous coworker capable of executing projects that span days rather than minutes. According to the company, Qwen3.8-Max can autonomously complete software projects lasting more than 10 days, reproduce research papers involving thousands of lines of code, perform iterative chip-design optimization, and continuously revise plans using multimodal feedback loops. Those demonstrations remain company-produced and have not yet been broadly replicated by independent evaluators. Nevertheless, they illustrate a growing industry trend: frontier models are increasingly competing on their ability to finish entire workflows rather than answer individual prompts. Benchmarks increasingly reward autonomous execution The benchmark suite released alongside Qwen3.8-Max reflects this shift. Instead of focusing solely on traditional reasoning exams or coding puzzles, many of the highlighted evaluations measure long-horizon execution. On OSWorld-Verified, which evaluates computer-use agents interacting with desktop environments, Qwen3.8-Max posts 86.1, ahead of GPT-5.6 Sol Max's 83.2, Fable 5's 85.0, and Gemini 3.1 Pro's 76.2. The model also leads: PaperBench: 93.0 TerminalBench 2.1: 86.6 Vision2Web: 69.0 LVBench: 81.8 ERQA: 77.8 Elsewhere, it remains competitive with proprietary leaders while trailing in several categories. On the professional software engineering benchmark SWE-Pro, for example, OpenAI's model posts the highest reported score, while Opus 4.8 continues to lead on certain software engineering evaluations and Agents' Last Exam. Rather than dominating every benchmark, Qwen appears to offer one of the broadest balanced performance profiles currently available. That balance may ultimately matter more for enterprise buyers than isolated benchmark wins. Many organizations increasingly evaluate models based on how reliably they complete heterogeneous workflows—writing code, reading documents, navigating interfaces, generating reports, inspecting images and coordinating multiple subtasks—rather than optimizing for one narrow capability. Where Qwen3.8-Max appears strongest Assuming Alibaba's published results translate into production deployments, several enterprise workloads stand out as particularly well suited for Qwen3.8-Max. 1. Long-running software engineering Alibaba's primary demonstration involves autonomous software development extending beyond ten days. While enterprises should treat these demonstrations as vendor claims until independently reproduced, they align with a growing interest in persistent coding agents that operate continuously rather than interactively. Organizations experimenting with autonomous engineering teams, CI/CD automation, repository maintenance, regression testing or feature implementation may find Qwen particularly attractive if its agentic performance proves consistent outside laboratory settings. 2. Computer-use agents The strongest differentiator may be computer use. OSWorld has rapidly become one of the industry's most closely watched benchmarks because it measures a model's ability to interact with operating systems instead of simply generating text. Models capable of reliably navigating desktop software can automate countless repetitive business processes, including document processing, enterprise software integration, internal operations and legacy workflows where APIs may not exist. Leading OSWorld could therefore translate into real operational advantages if benchmark performance generalizes to production environments. 3. Research automation Qwen's PaperBench leadership suggests strong potential for organizations performing scientific computing, literature review, experiment reproduction and technical analysis. Research institutions, pharmaceutical companies and industrial R&D teams increasingly use LLMs not only for summarization but also for executing reproducible computational workflows. Models capable of maintaining context across extended sessions become increasingly valuable in these environments. 4. Multimodal industrial workflows Unlike earlier multimodal systems that primarily analyze uploaded images, Qwen describes vision as an ongoing feedback mechanism integrated into planning and execution. That architecture could prove particularly useful in manufacturing, logistics, engineering inspection and design review, where visual inputs continuously inform operational decisions rather than serving as isolated prompts. The economics may prove just as important Perhaps the biggest competitive pressure comes not from benchmark scores but from pricing through Qwen's application programming interface (API) on QwenCloud (based in China): Qwen3.8-Max launches at $2/$6 per million input/output tokens, a mid-priced model but undercutting the top U.S. proprietary offerings to which it is benchmarked against by meaningful percentages, less than 1/3 the combined in/out price of Claude Opus 5 and less than 1/4 the price of GPT-5.6 Sol Max. Model Input ($/1M) Output ($/1M) Total ($/1M) Source MiMo-V2.5 Flash $0.10 $0.30 $0.40 Xiaomi deepseek-v4-flash $0.14 $0.28 $0.42 DeepSeek deepseek-v4-pro $0.435 $0.87 $1.305 DeepSeek GPT-5.6 Luna $0.20 $1.20 $1.40 OpenAI MiniMax-M3 $0.30 $1.20 $1.50 MiniMax LongCat-2.0 — limited-time promo $0.30 $1.20 $1.50 LongCat Gemini 3.1 Flash-Lite $0.25 $1.50 $1.75 Google Qwen3.7-Plus $0.40 $1.60 $2.00 Alibaba Cloud MiMo-V2.5 $0.40 $2.00 $2.40 Xiaomi Gemini 3.5 Flash-Lite $0.30 $2.50 $2.80 Google LongCat-2.0 — standard $0.75 $2.95 $3.70 LongCat MiMo-V2.5 Pro (≤256K) $1.00 $3.00 $4.00 Xiaomi GLM-5.2 $1.40 $4.40 $5.80 Z.ai Grok 4.5 $2.00 $6.00 $8.00 xAI MiMo-V2.5 Pro (>256K) $2.00 $6.00 $8.00 Xiaomi Qwen3.8-Max $2.00 $6.00 $8.00 QwenCloud Gemini 3.6 Flash $1.50 $7.50 $9.00 Google Qwen3.7-Max $2.50 $7.50 $10.00 Alibaba Cloud Gemini 3.5 Flash $1.50 $9.00 $10.50 Google Gemini 3.1 Pro Preview (≤200K) $2.00 $12.00 $14.00 Google GPT-5.6 Terra $2.00 $12.00 $14.00 OpenAI GPT-5.4 $2.50 $15.00 $17.50 OpenAI Kimi K3 $3.00 $15.00 $18.00 Moonshot AI Gemini 3.1 Pro Preview (>200K) $4.00 $18.00 $22.00 Google Claude Opus 5 $5.00 $25.00 $30.00 Anthropic GPT-5.5 $5.00 $30.00 $35.00 OpenAI GPT-5.5 Instant (chat-latest) $5.00 $30.00 $35.00 OpenAI Sakana Fugu Ultra (≤272K) $5.00 $30.00 $35.00 Sakana AI GPT-5.6 Sol — Standard mode $5.00 $30.00 $35.00 OpenAI Claude Fable 5 / Claude Mythos 5 $10.00 $50.00 $60.00 Anthropic GPT-5.6 Sol — Fast mode $10.00 $60.00 $70.00 OpenAI Lower inference costs increasingly matter because agentic systems consume dramatically more tokens than conventional chatbots — a reality that likely factored into OpenAI's decision late last week to cut the API prices of its mid- and lower-end GPT-5.6 lineup of models (Terra and Luna) by 20% and 80%, respectively. Indeed, as those running these systems can attest, multi-hour autonomous workflows, iterative planning and continuous self-correction can generate millions of tokens during a single task. For enterprises deploying hundreds or thousands of agents simultaneously, inference costs often become one of the largest operational expenses. Small reductions in per-token pricing therefore compound rapidly. How it compares with American frontier models Despite headline benchmark comparisons, Qwen3.8-Max should not necessarily be viewed as a wholesale replacement for leading American models. Instead, its strengths suggest different deployment strategies. OpenAI's GPT family continues to excel as a broadly capable enterprise reasoning platform with mature tooling, ecosystem integration and extensive commercial deployment. Organizations already invested in Microsoft ecosystems or OpenAI's enterprise offerings may continue to value those operational advantages even if Qwen leads on selected agent benchmarks. Anthropic's Claude Opus remains widely regarded as one of the strongest coding assistants, particularly for careful software engineering and long-context reasoning. Some enterprises may still prefer Claude for human-in-the-loop development where reliability and predictable behavior outweigh raw autonomy. Google Gemini continues to differentiate itself through deep Workspace integration, multimodal capabilities and Google Cloud services, making it attractive for organizations already standardized on Google's enterprise stack. Where Qwen appears most compelling is for enterprises prioritizing autonomous execution, extended planning horizons and favorable inference economics without sacrificing frontier-level performance. The open-weight question remains unanswered The largest unknown surrounding Qwen3.8-Max has little to do with benchmarks. Alibaba says open weights are coming next week. However, neither the announcement nor the provided documentation specifies the license that will govern those weights. That distinction could prove critical. A permissive license such as Apache 2.0 would significantly broaden enterprise adoption by allowing organizations to self-host, fine-tune and integrate the model into proprietary products with relatively few restrictions. A custom license—similar to approaches used by several recent frontier releases—could impose limitations on commercial deployment, redistribution, field of use or model modification. Such restrictions would narrow the appeal for enterprises seeking long-term infrastructure investments, regardless of the model's technical performance. Moonshot AI's recent Kimi K3 release illustrates why this distinction matters. While Kimi K3 made its weights openly available to all, its licensing terms included specific terms including a disclosure and a commercial license requirement for those offering it as a "Model as a Service." Until Alibaba publishes Qwen3.8-Max's license, organizations considering self-hosting should treat the open-weight announcement as promising but incomplete. An increasingly crowded frontier Qwen3.8-Max arrives during one of the fastest-moving periods in the history of foundation models. Within weeks, developers have seen major releases from Moonshot AI, OpenAI, Anthropic and others, each emphasizing different strengths: reasoning, coding, multimodality, autonomous agents or economics. Alibaba's contribution is notable because it combines competitive benchmark performance, aggressive pricing, a million-token context window and a stated commitment to releasing weights for its flagship model. Whether it becomes the preferred platform for enterprise autonomous agents will ultimately depend less on leaderboard positions than on broader independent validation, production reliability and the licensing terms accompanying the forthcoming weight release. Those factors—not benchmark charts alone—will determine whether Qwen3.8-Max becomes a genuine alternative to the leading American proprietary models or simply another impressive entrant in an increasingly crowded frontier AI race.

VentureBeat·August 3, 2026·9 min read
Asana's AI agents share memory across your company — but not your secrets
AITutorial

Asana's AI agents share memory across your company — but not your secrets

Enterprise teams building AI agents keep hitting the same wall: a chatbot that can answer a prompt but can't remember what the last five people asked it, and can't tell you whether last month's version actually worked. In a fireside chat with VentureBeat's Sam Witteveen at VB Transform 2026 , Asana's chief product officer, Arnab Bose, unpacked how his team tackled this problem to build a new operating system: Agentic Work Management (AWM). The product treats AI agents as coachable teammates that operate alongside humans rather than as one-to-one assistants. For product builders and developers trying to move beyond basic integrations, Bose provided a look under the hood. He detailed how Asana engineered AWM, offering a blueprint for solving real-world bottlenecks and building agentic systems at scale. The Work Graph: 18 years of company data, repurposed To build an operating system for human-agent teams, Asana needed a ready-made enterprise context graph. They built AWM on top of their 18-year-old architecture: the Work Graph. This graph-based database organizes information through a structure the company calls the Pyramid of Clarity. The smallest unit of work is a task with an assignee and a due date. Tasks belong to projects, projects roll up into portfolios, and portfolios connect to company-wide goals. The graph can help trace for example how a delayed design task impacts a corporate revenue goal. The Work Graph provides a real-time ledger of who does what, by when, and why. AWM leverages this architecture to create a multiplayer teammate. A standard AI copilot is stateless and tied to a single user's prompt. Because AWM plugs into the Work Graph, the AI can view overarching company goals, update project statuses, and share memory with human colleagues. "Because [the agent] is plugged into the Work Graph, it's not just looking at a particular prompt that you're sending it or looking at a particular individual's markdown file system on their local file,” Bose said. “It's working off of that shared ledger for the whole company." AWM is already in production. Bose said Asana has "several customers live and successful on it," including FedEx, which published its own case study on the shift. Building in guardrails for confidential work Shipping AWM to enterprise customers required Asana to solve several technical hurdles. The first was data governance. If an AI teammate acts across a company, it builds a shared memory by learning from workflows and human feedback. Bose highlighted a critical boundary problem: If an executive uses AWM to build workflows for a confidential project, the system must ensure the agent's updated memory does not leak context to an unauthorized employee who interacts with the same agent later. "[I] shouldn't be able to leverage that shared memory when I run the AI teammate if you created that memory using that same teammate on a project that is, let's say, a secret M&A project that I don't have access to," Bose said. Asana engineered a system of access controls to govern what triggers the creation of a memory versus the simple execution of a task. Second, AWM handles dynamic model routing to abstract prompt engineering away from the user. When a user assigns a task to an AI teammate (i.e., drafting a job description for a general manager role), the AI cross-references public job postings, Asana’s internal style guide, and product requirement documents. For a complex task, the system automatically routes the prompt to a heavy frontier model — Bose pointed to Anthropic's Opus and OpenAI's models as examples — while lighter tasks get down-leveled to something faster and cheaper. "We don't want the knowledge worker to have to think through what the best possible prompt, context engineering, and attachments are that they should put into the task," Bose said. "It should feel as if you were assigning the task to a human being." This dynamic routing introduces a third challenge: billing abstraction. Agentic tasks vary in computational complexity, making credit burn rates unpredictable. "We don't want to get into a state where our customers are having to reason about the fact that some of these tasks... are way more complex than others and they'll be burning credits at different rates," Bose said, adding that unpredictable pricing risked customers throttling their own employees by capping how often they could run an AI teammate. To make AWM commercially viable, Asana designed its billing architecture to charge a static cost per task completion. The platform absorbs the complexity of model selection, token counts, and run limits to ensure predictable enterprise pricing. The problem with stateless chatbots AWM targets a specific problem with current enterprise AI deployments: statelessness. Developers can easily connect large language models to enterprise tools like Slack, Google Drive, or Databricks using Model Context Protocol (MCP) integrations. However, basic chat-based agents lack persistence. Bose detailed a scenario where a user asks a chat agent to draft a marketing campaign based on historical performance and competitive research. The agent fetches data from external tools to answer the prompt, but the execution happens in a vacuum. It is a one-off task that benefits a single individual. It fails to create a reusable workflow for the next person building a similar campaign. "The challenge with that is that those calls are stateless, and they are not leveraging a shared company brain that is this graph-based database or a context graph," Bose said. AWM solves this by creating a permanent state. When an AI teammate inside AWM completes a task, the system records the metadata. It registers whether the completion improved the project status and how it moved higher-level company goals. Inside CoreWeave's product launches Cloud provider CoreWeave is an early adopter using AWM to overhaul complex new product launches. "CoreWeave is using both our deterministic AI studio workflow rules as well as multiple AI teammates to do new product launches," Bose shared. In the past, CoreWeave product managers filled out complicated forms detailing infrastructure, parameters, and costs. Human reviewers manually evaluated these forms and broke them out into specific tasks for finance, marketing, and hardware teams. Under the AWM workflow, a product manager writes a standard Google document pointing to their product requirement documents. A deterministic AI workflow reads the document, automatically creates the project structure, and assigns tasks. Specialized agents then take over the execution. One agent then watches overall project status and flags bottlenecks; another, working inside individual tasks, forecasts infrastructure costs and recommends approvals when the numbers align with historical budgets. The system automatically triages the busywork while human beings focus on evaluating the AI's outputs. The frenemy problem The dynamic gets complicated by the fact that the same frontier-model providers powering AWM under the hood — Anthropic, OpenAI — are also shipping their own competing agent products, like Anthropic's Claude in Slack (Tag). Pressed on the overlap, Bose didn't dispute the tension. " I think that's the reality that we all have to live in," he said. His case for AWM's staying power rests on Asana's 18 years of user-experience and workflow data, and prebuilt standard operating procedures for specific industries — expertise he argues raw frontier models don't have. A product like Tag can work well in Slack, he said, but it requires a highly curated channel and its own separate credentials for every downstream app it touches. "There's a big difference between the power of the model plus a lightweight way to demonstrate its value, and something that's pre-built … for true end-to-end use," Bose said.

VentureBeat·August 3, 2026·6 min read
How NTT DATA AIVista closes the last mile of agentic AI for enterprise agents
AIOpen Source

How NTT DATA AIVista closes the last mile of agentic AI for enterprise agents

Presented by NTT DATA AIVista At VB Transform 2026 , NTT DATA AIVista CEO Bratin Saha joined VentureBeat CEO and editor-in-chief Matt Marshall to discuss the last-mile challenge of operationalizing frontier models in regulated production, where reliability, context, guardrails, and security determine whether AI delivers enterprise value. The conversation centered around the question facing every enterprise now pouring money into AI: how to convert that spending into real, tangible value. "It's not just a model, you're building a system around the model," Saha said. The last mile is the work of wrapping a frontier model in an enterprise's own data, workflows, and guardrails. In the end, regulated production turns on more than just technology, Saha said. Today, most enterprise AI projects fail during implementation because of poor integration, domain specialization gaps, lack of governance, and unclear ownership of outcomes. Last-mile specialization turns a capable foundation model into an enterprise agent shaped by domain-specific workflows, risk appetite, client classifications, regulatory interpretations, and institutional knowledge. Why frontier models stall in enterprise workflows Frontier models fall well short of production-grade accuracy on many real-world insurance workflows, Saha said, but last-mile specialization can lift them to the reliability enterprises need. Out of the box, those models struggle with the complexity of regulated workflows such as multinational insurance claims. "These forms are pretty complex, often have handwriting, lots of checkboxes, and so on," he said, and that complexity is why frontier models like Fable 5, Opus 4.8, and GPT-5.5 fall short out of the box. Saha said the biggest gains come from specializing the entire AI system, not just the foundation model. That system gets specialized with the customer's data, workflow and, in many cases, the tribal knowledge that never made it into an operating procedure document. "The biggest bang for the buck comes from the specialization and then these specialized guardrails," he said. The work has three components: capturing the enterprise’s context and making it consumable by AI running an ensemble of models so cost does not go through the roof and adding specialized guardrails that check the model and force a redo when it gets something wrong. What the last mile of agentic AI actually requires None of this involves fine-tuning. VentureBeat’s latest enterprise survey found it ranked last among companies’ model-selection priorities. Instead, the last mile centers on domain knowledge and undocumented workflows that companies would never expose publicly without losing their competitive edge. "The last mile is about taking data that's proprietary to you and using that to build a system around the model that can steer the model in the right way that can put the appropriate guardrails around it," Saha said. In the end, enterprise AI is about moving a workflow from point A to point B rather than deploying a technology, and NTT's advantage comes from pairing AI experts with subject domain experts. "The only reason is because we go and talk to those human workers and we say, 'How do you actually do the work,'" he said. That expertise is then encoded into an agent. Success in insurance, manufacturing, and other regulated industries relies on three things at once, he added. "You need technology, you need the domain expertise, and you need the change management expertise," he explained, adding that across his team's clients, technology is not the bottleneck. How enterprises turn AI investment into tangible value For enterprises weighing large AI budgets, Saha's said the payoff comes not from the model but from the work built around it. "When you're deploying AI in the enterprise, you're not deploying a technology," he said. "You are taking a workflow that exists and taking it from point A to point B." The value is created by the workflow that gets moved, not the model that helps move it. That reorders where money should go. "Technology is not the bottleneck," Saha said, pointing instead to the domain expertise and change management wrapped around the model, and to the discipline of commiting to all three together. Spending aimed only at the model leaves most of the return on the table. Enterprises don’t have to choose between embedding AI into existing workflows and redesigning those workflows from scratch. NTT sees the two as successive stages of the same journey. "We are starting with embedding in the workflow because it's easier change management," he said, noting that customers running mission-critical operations will not let a vendor rip out a working process midstream. "Once that happens, then we go into, how can we now reimagine this? And that really is where the biggest bang is." Where enterprise AI stays bespoke and where it becomes scalable Keeping intelligence in the surrounding system rather than the model also preserves swappability and lets enterprises take advantage of open-weight and open-source models as they mature. Saha’s team runs an ensemble that mixes frontier and open-source models, and he expects the industry to lean on open weights wherever the cost of a mistake is low while reserving frontier reasoning for the cases that demand it. "In many situations, especially in regulated industries where mistakes are very expensive, that last extra couple of percent matters," he said. The platform follows the same pattern: Guardrail generation and neurosymbolic models scale across customers, while capturing each organization’s tribal knowledge remains bespoke. Saha pointed to NTT DATA’s position as one of the world’s largest insurance third-party administrators as an advantage in acquiring that expertise. "The ability to take that knowledge and trust that has been built over 20 years is very hard to replicate instantly, and I do think that is a durable aspect of what we have," he said. Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com .

VentureBeat·August 3, 2026·5 min read
Stop graphing everything: When GraphRAG actually beats vector RAG
AIResearch

Stop graphing everything: When GraphRAG actually beats vector RAG

If you have built anything with retrieval-augmented generation (RAG) in the last two years, you have lived its central frustration: You chop your documents into chunks, embed them, retrieve the top few that look similar to the question, and hand them to the model. For “What was our Q3 refund policy?” This works beautifully. For “What are the recurring themes across two years of customer complaints?” it falls flat — because no single chunk contains the answer. The fashionable fix is GraphRAG : Instead of feeding the model isolated snippets, you first build a knowledge graph of the entities and relationships in your corpus, then use that structure as context. The pitch is seductive. But seductive pitches deserve scrutiny, so I went through the evidence — the original Microsoft paper plus four independent benchmark studies — to answer a simple question: When you swap text chunks for a context graph, do answers actually get better? The short version: Yes, substantially — but only for the right kind of question, and not for free. Let me show you the receipts. Why text chunks hit a wall Standard vector RAG retrieves the k passages most similar to your query. That design has three structural blind spots: It can’t connect the dots. When an answer requires joining facts that live in different passages through a shared entity, chunks embedded in isolation never reveal the link. It’s blind to global questions. “What are the main themes?” needs the whole corpus, but similarity search only returns the handful of chunks that superficially resemble the question. It severs context at chunk boundaries. The relationships and hierarchy that complex reasoning depends on are exactly what chunking throws away. Microsoft Research framed this crisply when they introduced GraphRAG: Baseline RAG “struggles to connect the dots” and performs poorly when asked to “holistically understand summarized semantic concepts over large data collections.” What a context graph changes GraphRAG attacks the problem before any question is asked. During indexing, a large language model (LLM) reads every chunk and extracts entities, relationships, and claims, assembling them into a weighted knowledge graph. It then runs community detection (the Leiden algorithm ) to cluster the graph into a hierarchy of related topics, and pre-writes a natural-language summary for each community. At query time, those summaries do the heavy lifting. Each relevant community drafts a partial answer (the “map” step), the partials are ranked and merged (the “reduce” step), and the model synthesizes a final response grounded in structure rather than in a few cherry-picked snippets. Variants like HippoRAG take a different route, using the graph plus a Personalized PageRank walk to find the right passages — but the core idea is the same: Let relationships, not just cosine similarity, decide what context the model sees. The evidence: Four studies, one pattern 1. Global sense making: The headline win Microsoft pitted GraphRAG head-to-head against naïve RAG on global, “make sense of the whole corpus” questions over million-token datasets, with an LLM acting as judge across three axes: Comprehensiveness, diversity, and empowerment. GraphRAG won 72 to 83% of comprehensiveness comparisons and 62 to 82% of diversity comparisons against vector RAG . Its highest-level summaries used up to 97% fewer tokens than processing the source text directly. That is not a rounding-error improvement. On exactly the kind of question that breaks text-chunk RAG, the graph wins two out of three times or better. 2. Multi-hop retrieval: The graph finds what chunks miss The second piece of evidence is about retrieval quality: Does the right supporting passage even make it into the top results? On the standard multi-hop QA benchmarks (MuSiQue, HotpotQA, 2WikiMultiHopQA), graph-guided retrieval lifts Recall@5 dramatically: Average Recall@5 climbs from 73.4% (naïve RAG) to 87.8% (graph-guided), a +19.6 point gain. The biggest jumps come on the hardest, cross-document sets: +31 points on MuSiQue and +28 points on 2Wiki . HippoRAG reports up to a 20% accuracy improvement on multi-hop QA, at 10–20× lower cost and 6–13× faster than iterative retrieval methods. 3. The controlled head-to-head - where it gets honest Here is where the story gains nuance. A 2025 study from Michigan State and Meta ran RAG against four GraphRAG families under one unified protocol — identical chunking, embeddings, and generation — and found no single winner. The two approaches are complementary: On single-hop, factual lookup (natural questions), plain RAG edged ahead (F1 64.8 vs. 63.0 for the best graph method). On multi-hop reasoning (MultiHop-RAG), graph-guided retrieval pulled in front (70.3 vs. 67.0 overall accuracy). The lesson: A context graph is not a universal upgrade. It is a specialized one that pays off precisely when questions demand reasoning across pieces. 4. When to use graphs: The task-type verdict The most recent benchmark, GraphRAG-Bench (ICLR 2026), set out to answer “In which scenarios do graph structures provide measurable benefits?” Its accuracy-by-task numbers map the boundary cleanly: Simple fact retrieval: Text chunks 60.9 vs. graph 60.1 — effectively a tie. The graph’s structure is overhead the query doesn’t need. Complex reasoning: Graph 53.4 vs. chunks 42.9 — a +10 point graph win. Contextual summarization: Graph 64.4 vs. chunks 51.3 — a +13 point graph win. The scorecard Read top to bottom, the pattern is unmistakable: The graph’s advantage grows with the reasoning depth of the question, while text chunks hold their ground on isolated facts. The catch: Cost and the LLM-judge problem Two caveats keep this from being a slam dunk, and ignoring them is how teams end up disappointed. Building the graph is expensive. Having an LLM extract entities and relationships from an entire corpus isn’t cheap. One analysis put index construction at roughly $48 against GPT-4o for a moderate corpus, far above a vanilla vector index. (Microsoft’s own follow-up, LazyGraphRAG, defers extraction to query time and cuts that to around 0.1% of the cost - a tacit admission that the original budget is impractical for many deployments.) Many of the wins are judged by another LLM — and LLM judges are biased. An independent audit found systematic flaws in this evaluation style: position bias (swapping which answer appears first can swing the win-rate by more than 30 points), length bias , and trial bias (identical comparisons disagree across runs). After correction, one popular method’s reported 66.7% win rate fell to about 39% — below the 50% break-even line. The takeaway is not “the research is wrong.” It is that the large gains — the +20% multi-hop accuracy, the +15-to-30-point recall jumps — are robust, while narrow comprehensiveness margins deserve a skeptical second look with reference-based metrics. So when should you reach for a context graph? Strip away the hype and the decision is refreshingly practical. Use a context graph when: Your questions are multi-hop, global, or sensemaking in nature; you need comprehensive, multi-perspective answers; and your corpus is richly interconnected (research libraries, case files, incident histories, knowledge bases). Stick with text chunks when: Your queries are mostly single-fact lookups; your corpus is small or flat; and indexing cost, latency, and operational simplicity outweigh a marginal quality bump. Best of all, go hybrid: The systematic studies converge on the same recommendation: route each query to the right method, or fuse evidence from both. Combining graph and chunk retrieval consistently beats either one alone. You don’t have to choose a religion; you have to build a router. The bottom line A context graph is not magic, and it is not snake oil. It is a targeted instrument. Hand it a question that requires connecting scattered facts or synthesizing a whole corpus, and it will outperform text chunks decisively. Hand it “what’s the phone number on page 3,” and you’ve paid for indexing you didn’t need. The teams that win with GraphRAG in 2026 won’t be the ones who graph everything. They’ll be the ones who know which questions deserve a graph — and build pipelines smart enough to tell the difference. Dattaraj Rao is an R&D architect at Persistent Systems

VentureBeat·August 2, 2026·7 min read
Structured AI data pipelines score 10.9 points below free-form code — DataFlow-Harness closes the gap
AIOpen Source

Structured AI data pipelines score 10.9 points below free-form code — DataFlow-Harness closes the gap

If you ask an AI coding agent to write a standalone Python script to parse a single JSON file, it will likely give you a perfect answer in seconds. But the same agent often breaks if you ask it to build a systematic data processing pipeline, like ingesting thousands of messy documents, chunking text, scoring quality, and filtering noise for a Retrieval-Augmented Generation (RAG) system that fits your specific enterprise stack. While large language models (LLMs) excel at one-off code generation, their outputs for complex data-processing tasks are typically free-form, disposable scripts. These scripts are detached from the governable workflow abstractions that MLOps teams rely on for production, making them difficult to audit or edit visually. To address this, researchers at Peking University, Zhongguancun Academy, and Shanghai’s Institute for Advanced Algorithms Research introduced DataFlow-Harness , an open-source framework that guides an LLM agent to build structured, visual data-processing workflows step-by-step, rather than writing raw code from scratch. The framework makes AI-generated pipelines easier to manage and integrate into existing architectures because the generated artifacts are persistent and easily editable. The researchers report that the platform achieves a 93.3% observed end-to-end pass rate on a 12-task data-engineering benchmark. Compared to standard Claude Code, it reduces API costs by up to 72.5% and response latency by 49.9%, while achieving nearly the same success rate as an AI given the entire codebase to write standard scripts. For enterprise teams, this means getting the speed of AI automation without accumulating unmanageable technical debt, ensuring that pipelines remain secure, auditable, and ready for production. The "NL2Pipeline gap" Data-centric AI requires workflows for tasks like synthetic data generation, retrieval augmentation, and model training. While LLMs can translate natural language into executable implementations to perform these tasks, high task accuracy is insufficient for production deployment. "The first wall is usually not writing Python," Runming He, first author of the DataFlow-Harness paper, told VentureBeat. "Modern coding agents can often produce a plausible script quickly. The harder problem is grounding that script in a live production platform: using operators that are actually installed, matching the real dataset schema, referring to registered datasets and model services, preserving dependencies between stages, and leaving behind an artifact that another engineer can understand and revise." General-purpose AI agents frequently hallucinate dependencies, relying on unavailable operators or outdated platform assumptions. Instead of leaving behind an artifact that another engineer can understand and revise, they generate disposable code that is difficult to audit through workflow managing tools. The researchers define this challenge as the "NL2Pipeline gap": the disconnect between a user expressing workflow requirements in natural language and the production environment requiring structured and persistent pipeline assets. The researchers demonstrated this gap in their experiments. For example, when Claude Code was allowed to write standard, free-form scripts using codebase context, it hit a 94.2% success rate. However, when restricted to only using the platform's specific building blocks to create a native workflow graph, its success rate dropped to 83.3%. This gap is the paper's central finding: native, governable pipelines are meaningfully harder for the agent to produce than throwaway code. “Closing this gap requires more than improving code-generation accuracy: construction must remain grounded in platform semantics and produce artifacts that integrate with the host platform,” the researchers write. How the four components work together "DataFlow-Harness changes the agent’s action space," He said. "Instead of asking the agent to emit arbitrary code, it retrieves the live operator registry and current pipeline state through MCP and applies typed, incremental changes to a persistent DAG." To achieve this, the platform organizes workflow synthesis around four components: the Data Pipeline Backend, the interaction layer (DataFlow-WebUI), the MCP Tools Layer, and the AI guidance layer (DataFlow-Skills). The Data Pipeline Backend acts as the authoritative source of truth across conversational, visual, and programmatic interfaces. It represents the pipeline as a directed acyclic graph (DAG), a structured workflow map containing data sources, configured pre-built processing modules (which the researchers refer to as "operators"), and execution dependencies. Instead of generating free-form code, agents interact with this backend through “typed mutations,” like adding an operator or connecting edges. DataFlow-Skills are markdown files that inject domain-specific knowledge into the model's context window, guiding it on operator-selection patterns, schema inference, and assembly procedures. Rather than letting the AI guess how to assemble components, skills provide the AI with compatibility rules, teaching it how to correctly match different data formats and handle complex data structures without breaking the pipeline. The MCP Tools Layer gives the AI access to the operator registry and current state of the data workflow. The AI proposes structured changes through the tools layer. The system validates the changes to ensure the workflow runs in a valid sequence and that every connected module speaks the same data language. DataFlow-WebUI provides two interfaces that allow humans and AI to build the workflow together. Developers can describe workflow requirements in natural language through a conversational interface. They can also access the workflow as a graphical map in a visual DAG editor. Here, they can directly inspect the changes proposed by the AI and make modifications. “The current implementation performs static checks against platform metadata before accepting pipeline changes,” He said. “These include checks for registered datasets, operators and model-serving references, field flow, and some invalid parameter usage, as well as structural validity. The result is visible in a graphical editor and can be revised either manually or by the agent in later turns.” The results: 93.3% pass rate, 72.5% lower cost The researchers tested DataFlow-Harness on a benchmark of 12 tasks across six industrial data-processing scenarios, such as QA generation, review governance, and schema normalization. They used Claude Opus 4.7 as the backbone model in their experiments. They compared DataFlow-Harness against three baselines: Vanilla CC: An unconstrained coding baseline using standard Claude Code. Context-Aware CC: An agent that has access to the DataFlow codebase in its context window. MCP-only: An agent that has access to the DataFlow MCP tools and is instructed to generate platform-native DAGs (without access to DataFlow-Skills). DataFlow-Harness achieved a 93.3% end-to-end pass rate, improving by 10.0 percentage points over MCP-only and beating Vanilla CC (91.7%), while being within 0.9 percentage points of Context-Aware CC (94.2%). Importantly, it reduced API costs to $0.261 per task, a 72.5% drop compared to Vanilla CC and 42.8% compared to Context-Aware CC. In generating workflows, it was 49.9% faster than Vanilla CC and 17.6% faster than Context-Aware CC. DataFlow-Harness proved particularly effective on complex tasks that depend on implicit domain knowledge, like QA generation. The baseline MCP-only approach frequently generated structurally valid DAGs but struggled to infer task-specific procedures from operator descriptions alone. To show how this works in the real world, the researchers detailed a textbook-to-VQA extraction task. This job required the AI to stitch together capabilities such as PDF parsing, layout recovery, OCR, figure extraction, multimodal understanding, and long-range question-answer matching. DataFlow-Harness achieved 97.2% precision and an 87.3% coverage rate, easily beating the baselines. By having the AI snap together existing platform assets rather than coding complex tasks from scratch, it recovered more valid QA pairs from the document. Their experiments also showed that DataFlow-Harness is highly effective at creating data generation pipelines. For example, in a synthetic instruction-data generation task, the agent built a multi-stage pipeline that generated candidate instruction–response pairs, critiqued and rewrote them, scored them with an LLM-based judge, and filtered low-quality outputs before training. "Such workflows are costly to build and fragile to maintain as collections of ad hoc scripts," He said. "The harness does not make them automatically safe, but it turns them into explicit, editable stages that engineers can inspect, test, and govern using normal production controls." Similarly, when tasked with building a math data cleaning-and-synthesis pipeline, the data produced by the DataFlow-Harness pipeline trained a better-performing model with higher average accuracy on AIME24 and AIME25 benchmarks than the data produced by the vanilla Claude Code pipeline. Tech stack fit and implementation tradeoffs For engineering teams evaluating DataFlow-Harness, it is important to understand how it fits into existing infrastructure. Released under the Apache 2.0 license, the current implementation requires a bit of engineering to fit into popular tech stacks. "The current implementation is native to the DataFlow platform; it is not a turnkey Airflow, Prefect, or Spark plug-in," He said. To use those systems as an execution backbone, teams must build an adapter to connect their organization’s registry, metadata, and execution interfaces to the agent's control layer. Furthermore, organizations must invest in the boundaries they want the AI to respect. This requires maintaining an operator registry, defining schemas, and encoding recurring domain procedures as Skills. Because of this overhead, He recommends against using the framework for small, one-off transformations where a simple script suffices, or in legacy environments that cannot expose reliable metadata. Finally, while the platform prevents illogical connections by validating structural properties, it is an engineering control layer, not a compliance substitute. "The harness should still be treated as an engineering control layer, not as a substitute for compliance policy, validated detection models, access controls, audit logging, or human approval," He said. The platform is open-source, and developers can access the source code and codebase documentation directly via the project's GitHub repository . As protocols like MCP become standardized, the boundary between human engineers and AI agents will shift. "The goal is not autonomous data engineering without oversight," He said. "It is a better division of labor: agents perform repetitive construction inside explicit boundaries, while engineers remain responsible for the semantics, policies, and consequential decisions that require domain accountability."

VentureBeat·July 31, 2026·8 min read
How is your enterprise tracking AI agent telemetry? Groundcover thinks it should never leave your cloud
AIResearch

How is your enterprise tracking AI agent telemetry? Groundcover thinks it should never leave your cloud

The AI agent observability space is taking off — but how can enterprises be sure what observability products and solutions they need? Observability startup groudcover (lower case "g" intentional) announced this week that it raised $100 million in a round led by One Peak, bringing its total funding to $160 million. The company says it has more than 250 paying customers, tripled annual recurring revenue over the past year and is increasingly replacing established observability platforms inside enterprise environments. Those are company-reported figures, but together they point to growing momentum in one of enterprise software's most competitive markets. That market has long been dominated by companies including Datadog, Dynatrace, New Relic, Splunk and Grafana. Between them, they represent billions of dollars in annual revenue and years of product maturity. Breaking into that group has never been easy. groundcover's argument is that artificial intelligence has fundamentally changed the assumptions those platforms were built on. Rather than competing feature for feature, the four-year-old company is trying to convince enterprises that the architecture underpinning observability itself needs to change as AI systems become more autonomous, produce vastly more telemetry and increasingly participate in software operations. Whether that thesis proves correct remains an open question, but it offers a compelling lens through which to examine how observability is evolving alongside enterprise AI. AI is turning telemetry into an infrastructure problem Observability has traditionally been viewed as a post-production discipline. Engineers deploy applications, monitor logs, metrics and traces, investigate incidents, and improve reliability over time. That workflow is changing. AI-assisted software development has dramatically accelerated deployment cycles. Coding assistants generate more code, infrastructure evolves more rapidly, and organizations are deploying increasingly complex distributed systems that combine microservices, Kubernetes clusters, APIs and large language models. At the same time, enterprises are beginning to operate AI agents that execute multi-step workflows, call external tools and interact with production systems. Each of those activities generates telemetry. The result is an explosion of operational data that organizations increasingly want to retain rather than discard. AI applications introduce additional layers of observability beyond traditional infrastructure monitoring, including prompt execution, model latency, token consumption, retrieval pipelines, tool invocations and agent behavior. As enterprises experiment with autonomous systems, that telemetry becomes increasingly valuable because it provides the context needed to understand what an AI system actually did and why. For many organizations, this creates tension with pricing models that charge according to the amount of data ingested. Historically, engineers have often responded by sampling traces, shortening retention periods or limiting which data is collected. Those approaches reduce costs, but they also reduce visibility precisely when AI-driven systems demand more complete operational context. "We've seen telemetry exploding," groundcover co-founder and CEO Shahar Azulay said during a recent media briefing. "Users are frustrated by not getting all the value from Datadog and similar platforms. They're limiting the data, siloing it, sampling it." Whether that frustration is widespread enough to reshape the market remains to be seen, but the underlying trend is difficult to ignore. AI is making observability less about collecting enough data and more about collecting everything organizations may eventually need. Rather than adding AI, groundcover argues the architecture itself has to change Many observability vendors have introduced AI assistants, AI-powered root cause analysis and AI observability features over the past two years. Datadog, Dynatrace, New Relic and Grafana have all announced products aimed at helping enterprises monitor AI applications or automate operational tasks. groundcover acknowledges those developments but argues they do not address what it sees as the more fundamental issue: where telemetry lives and how customers pay for it. Instead of operating a conventional SaaS platform that stores customer telemetry in vendor-managed infrastructure, groundcover uses what it calls a bring-your-own-cloud (BYOC) architecture. Customers keep the data plane—including telemetry storage and processing— inside their own AWS, Microsoft Azure or Google Cloud environments, while groundcover provides a managed control plane and user experience. A fully self-hosted deployment option is also available. While some competitors, including Datadog and a few other observability vendors, do offer limited hybrid or customer-controlled data residency options, these are generally not equivalent to a full BYOC model. In most cases, telemetry is still processed and stored within the vendor’s managed infrastructure, with only partial controls (such as regional data residency, private links, or selective log forwarding) available. That architectural decision influences nearly every aspect of the company's strategy. Because customers already pay for their own cloud infrastructure, groundcover argues it can avoid charging based on telemetry ingestion. Instead, pricing is based primarily on monitored hosts, regardless of telemetry volume. The company believes this changes customer behavior. Rather than deciding which logs or traces are too expensive to keep, organizations can theoretically retain complete telemetry and use it for operational analysis, compliance and AI-assisted troubleshooting. "We don't price by data volume," Azulay said. "We price by the size of the infrastructure." The distinction matters because AI workloads tend to increase telemetry far faster than infrastructure itself. That does not necessarily make host-based pricing universally cheaper. Organizations with relatively light workloads spread across many hosts may find different economics than dense Kubernetes environments generating enormous amounts of telemetry. The company's own briefing notes that per-host pricing is most advantageous for organizations with high telemetry density and may be less compelling for lightly utilized fleets. Still, the broader argument is less about cost alone than predictability. Enterprise infrastructure teams often struggle with observability bills that fluctuate alongside application growth. groundcover's model attempts to align pricing more closely with infrastructure planning rather than data generation. eBPF sits at the center of the company's technical differentiation The second pillar of groundcover's strategy is eBPF, a Linux kernel technology that has rapidly become one of the most important building blocks for modern cloud observability. Instead of requiring developers to manually instrument applications, eBPF allows software running inside the operating system kernel to observe network traffic, system calls and application behavior with minimal code changes. That enables faster deployment and broader visibility across infrastructure. For organizations operating Kubernetes clusters and cloud-native applications, reducing instrumentation complexity can significantly shorten deployment times while increasing telemetry coverage. Azulay argues this becomes especially important as AI systems generate increasingly complex interactions across services. "Our sensor allows us to observe systems very deeply from infrastructure to application to AI workloads without developers needing to instrument code," he said during the briefing. eBPF itself is hardly unique. Many observability vendors now incorporate it into their platforms. What groundcover argues differentiates its approach is combining automatic eBPF collection with customer-controlled storage, OpenTelemetry compatibility and unified pricing inside a single platform. The company's own research briefing acknowledges that none of these technologies individually represents a competitive moat. The claimed differentiation lies in the combination of eBPF-first collection, managed BYOC architecture, host-based economics and full-stack observability delivered together. AI agents are becoming both customers—and users—of observability Perhaps the most interesting aspect of groundcover's strategy extends beyond traditional monitoring. The company increasingly describes observability as infrastructure for autonomous software development. Historically, observability platforms have served human operators investigating production incidents. groundcover believes future observability platforms will increasingly serve AI agents as well. Its Agent Mode product allows engineers to investigate incidents using natural language across logs, metrics, traces and Kubernetes events. More importantly, Azulay envisions observability becoming the feedback mechanism that informs coding agents about what actually happened in production. Rather than simply detecting failures after deployment, observability becomes continuous operational context that autonomous systems can use to evaluate changes, identify regressions and eventually recommend or implement fixes. "We're seeing observability moving from being a post-production tool... to people taking context from production and feeding it back to their coding agents so they can write code better," Azulay said. Today, the company emphasizes that humans remain in the loop. Agent Mode investigates incidents and surfaces recommendations, but production changes still require human approval. Azulay expects autonomy to increase gradually as organizations become more comfortable allowing AI systems to participate in operational workflows. That vision reflects a broader trend emerging across enterprise software, where AI agents increasingly span development, testing, deployment and operations rather than functioning as isolated assistants. Why some enterprises are considering alternatives groundcover is entering an intensely competitive market populated by vendors with decades of enterprise experience. Datadog alone generated more than $3 billion in annual revenue in 2025 . Dynatrace, Cisco's Splunk business, Grafana Labs and New Relic all maintain extensive partner ecosystems, mature integrations and enterprise support organizations that newer entrants cannot easily replicate. groundcover is not attempting to outscale those incumbents overnight. Instead, it argues that AI creates an architectural inflection point similar to previous transitions from on-premises infrastructure to cloud-native computing. According to Azulay, many customers initially adopt groundcover to reduce observability costs but increasingly remain because they want unrestricted access to richer telemetry and AI-native workflows. He says deployments typically replace incumbent platforms rather than operate alongside them, although the company has not publicly disclosed customer migration data or independent studies validating that claim. The company's journalist briefing also urges caution around some performance claims. Revenue growth, customer counts and enterprise adoption figures originate from groundcover itself. Published customer case studies reporting significant cost savings are vendor-authored and should not be treated as independent validation without additional evidence. The briefing also recommends scrutinizing exactly what metadata leaves customer environments in standard BYOC deployments, rather than assuming that no operational data ever reaches vendor infrastructure. Those caveats are important because the observability market has become crowded. Gartner currently tracks more than one hundred observability products, and nearly every major vendor now markets AI-powered operational capabilities. Success will likely depend less on whether AI matters—which increasingly appears inevitable—and more on whether enterprises conclude that existing architectures remain sufficient. The larger question investors are betting on Viewed narrowly, groundcover's Series C is another large infrastructure funding round. Viewed more broadly, it reflects a growing debate about what observability becomes in an era where software increasingly writes, tests and operates itself. If AI continues generating exponentially larger volumes of operational data, traditional assumptions about telemetry collection, pricing and storage may come under increasing pressure. Vendors that built businesses around charging for data ingestion may need to evolve their economics alongside customer expectations. New entrants, meanwhile, have an opportunity to design around those changing assumptions from the outset. groundcover believes that opportunity lies in combining customer-controlled infrastructure, automatic telemetry collection and AI-assisted operations into a platform designed for autonomous software rather than simply adding AI features to existing observability products. Whether that architectural bet proves durable will depend on enterprise adoption over the next several years. But the company's latest funding round suggests at least some investors believe the next battle in observability will not be fought over dashboards or alerts. It will be fought over who builds the operational data layer that increasingly intelligent software relies upon to understand—and eventually manage—the systems it runs.

VentureBeat·July 31, 2026·9 min read
Not just OpenAI: Now Anthropic says its internal models got online and cyberattacked 3 other organizations
AISecurity Advisory

Not just OpenAI: Now Anthropic says its internal models got online and cyberattacked 3 other organizations

Days after OpenAI disclosed that two frontier AI models escaped containment measures and autonomously cyberattacked the AI code sharing platform Hugging Face, OpenAI's top U.S. rival Anthropic tonight revealed that — lo and behold — it has also had models surreptitiously access the web when they weren't supposed to, and cyberattack and gain "unauthorized access" to three other organizations. Anthropic says that it ran "capture the flag" cybersecurity scenarios with three models — Claude Opus 4.7, Claude Mythos 5, and unnamed internal research prototype — with its partner, the AI security firm Irregular . Anthropic says the models were not supposed to have internet access, but that a misunderstanding with Irregular allowed them to access the internet. Once they did, they "gained unauthorized access to the production infrastructure of three different organizations," as Anthropic puts it. Anthropic's blog post continues: "Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment." Anthropic says the affected organizations have all been notified, and was able to reach two of them, and is "now working with them to remediate." The third hasn't yet been reached. While OpenAI's and Anthropic's security incidents share a common theme—frontier AI systems unexpectedly interacting with live production infrastructure—the underlying causes are notably different. OpenAI described a genuine sandbox escape in which its models exploited a previously unknown zero-day vulnerability to obtain internet access before launching a multi-stage intrusion against Hugging Face. Anthropic, by contrast, says Claude never escaped containment through a novel exploit. Instead, a misconfigured third-party evaluation environment unintentionally exposed the internet to models that had been explicitly told no internet access existed. The distinction matters. Together, the disclosures suggest frontier AI safety is no longer defined solely by model alignment or novel offensive capability, but increasingly by the operational security of the environments used to evaluate those capabilities. What Anthropic found Anthropic said it reviewed 141,006 cybersecurity evaluation runs after OpenAI published its Hugging Face report . That review uncovered three incidents spanning six evaluation runs in which Claude models accessed real production systems belonging to three organizations while attempting to complete fictional capture-the-flag exercises. The incidents involved Claude Opus 4.7, Claude Mythos 5 and an unreleased internal research model. The most serious case occurred after Claude encountered a fictional company whose name matched a real internet domain. Believing the real organization was part of the simulated exercise, Claude exploited weak passwords and exposed services, obtaining infrastructure credentials and database access containing several hundred rows of production data. A second incident was arguably even more unusual. Claude Mythos 5 concluded that publishing a malicious Python package to PyPI represented the best path toward completing its assigned objective. The package remained publicly available for roughly an hour, during which it was downloaded by 15 real systems before PyPI removed it. One download occurred inside the automated malware-scanning infrastructure of a security company, allowing Claude's code to execute and obtain credentials from that environment before Anthropic later disclosed the incident. The third incident involved an internal research model that scanned approximately 9,000 internet-facing systems before compromising one organization using well-known techniques including exposed debug credentials and SQL injection. Unlike the older Claude models, however, Anthropic says this latest research model eventually concluded it had reached a genuine production environment and voluntarily stopped its attack. Similar outcomes, different root causes Although the headlines inevitably invite comparison with OpenAI's disclosure, the two incidents illustrate different failure modes. OpenAI's investigation concluded that its models actively discovered and exploited a previously unknown vulnerability in an Artifactory package-cache proxy to escape an otherwise isolated evaluation environment before later compromising Hugging Face through additional chained vulnerabilities. The company characterized that event as an "unprecedented cyber incident" involving state-of-the-art offensive AI capabilities. Anthropic's report instead argues its models never needed to break containment because containment had already been weakened through operational misconfiguration. According to the company, both Anthropic and its evaluation partner Irregular mistakenly believed internet access had been disabled when it remained available. Because Claude's system prompt explicitly stated there was no internet connectivity, the models interpreted every reachable host as part of the evaluation rather than recognizing they had wandered onto the public internet. Anthropic therefore characterizes the incidents primarily as infrastructure and evaluation-harness failures rather than evidence of models independently pursuing unauthorized goals. The reports nevertheless converge on one uncomfortable conclusion: frontier AI systems are increasingly capable of executing long-horizon offensive cyber operations whenever evaluation environments permit them to do so. Four major enterprise security takeaways so far... For enterprise security leaders, Anthropic's disclosure arguably shifts the conversation beyond "Can frontier models escape?" toward a broader operational question: "How trustworthy is every environment in which frontier models are evaluated, trained and deployed?" There are at least 4 lessons to be learned: The first lesson is that evaluation infrastructure itself now deserves production-grade security engineering . Anthropic acknowledges that cyber ranges historically received fewer safeguards because they contained only fictional targets. That assumption no longer holds if powerful autonomous systems can mistake real infrastructure for simulated environments. Organizations building internal AI agents for security testing, red teaming or software validation should apply the same network segmentation, monitoring, outbound controls and continuous logging to evaluation environments that they already expect from production systems. Second, both disclosures reinforce that alignment alone cannot compensate for environmental ambiguity . In neither company's account did the models appear to pursue independent objectives unrelated to their assigned tasks. Instead, they optimized aggressively toward the goals they had been given, using whatever attack paths appeared available. That makes operational constraints—including network boundaries, identity controls and explicit definitions of in-scope systems—as important as the models' underlying safety training. Third, enterprises deploying increasingly autonomous AI agents shoul d treat situational awareness as a security dependency rather than an academic capability. Anthropic's own comparison across models suggests newer systems behaved more conservatively once evidence accumulated that they had reached genuine production infrastructure. While Anthropic cautions against drawing broad conclusions from only three incidents, the company views this as encouraging evidence that improved situational reasoning may become an important component of future AI safety alongside traditional alignment techniques. Finally, these two disclosures together mark an inflection point for enterprise threat modeling. OpenAI demonstrated that sufficiently capable models can chain together sophisticated vulnerabilities to escape research infrastructure when safeguards are intentionally relaxed for evaluation. Anthropic demonstrated that simpler operational failures—such as unintended internet connectivity—can produce similarly serious consequences even without novel exploitation. The common denominator is not any single vendor or model family. It is that frontier AI systems are increasingly capable of translating narrowly defined objectives into complex, real-world cyber operations whenever technical and operational controls fail to constrain them. For enterprise CISOs, that means AI safety can no longer be viewed solely as a model problem. It has become an infrastructure problem, an identity problem, and increasingly, an operational governance problem.

VentureBeat·July 31, 2026·6 min read
Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size
ProgrammingCertification

Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size

Just two weeks after Thinking Machines released Inkling , its first open source AI language model, the well-funded startup led by former OpenAI chief technology officer Mira Murati today introduced Inkling-Small without sacrificing much of any performance — and in fact, the new model surpasses its larger predecessor on several benchmarks. Inkling Small is a 276-billion-parameter multimodal reasoning model with a permissive Apache 2.0 license that comes within a single point of its larger sibling on the third-party Artificial Analysis Intelligence Index , despite the original Inkling being 975 billion parameters (internal model settings). It accepts text, image and audio inputs, produces text, and supports a context window of up to one million tokens. Inkling Small uses 12 billion active parameters per token, compared with Inkling’s 41 billion active parameters, while preserving much of the flagship’s coding, reasoning and multimodal performance. For enterprises, the appeal is not simply that Inkling-Small is smaller. It is that developers appear to give up relatively little capability while reducing the model’s compute requirements, inference costs and deployment footprint. The model remains far too large for a laptop or conventional workstation, but it is materially easier to operate than the 3.5X larger flagship, making it a good fit for enterprises with some — but not a lot — of their own graphics processing units (GPUs). Thinking Machines has released the full weights on Hugging Face and added support for fine-tuning through its Tinker model training application programming interface (API). At launch, the company is advertising a limited-time 50% discount, bringing API pricing for the standard 64K-context Inkling-Small model to $0.58 per million prefill (input) tokens, $1.44 per million sampled (output) tokens, and $1.73 per million training tokens, with cached prefill requests priced at $0.116 per million tokens. A 256K-context variant is also available at higher rates. Nearly the same performance at a quarter the size Artificial Analysis assigned Inkling-Small a score of 40 on its Intelligence Index, compared with 41 for Inkling. That result is notable because Inkling-Small has 276 billion total parameters and 12 billion active parameters, while Inkling has 975 billion total parameters and 41 billion active parameters. Artificial Analysis also reported that no open-weight model at Inkling-Small’s size or smaller scored higher on the index. The model does more than merely approach the flagship’s aggregate score. On several evaluations, it surpasses Inkling. Thinking Machines reports that Inkling-Small scores 80.2% on SWE-bench Verified, compared with Inkling’s 77.6%, and 64.7% on Terminal Bench 2.1, compared with 63.8% for the larger model. It also edges ahead on SciCode, Humanity’s Last Exam, GPQA Diamond and CritPt. The gains are not universal. Inkling retains a clear advantage on factual knowledge and some agentic tasks. Inkling-Small scores 15.5% on τ³-Banking, compared with 23.7% for Inkling, and its AA Omniscience score is negative, reflecting weaker factual coverage even though its reported hallucination rate is slightly lower. That tradeoff matters for enterprises. Inkling-Small may be attractive for coding assistants, tool-use systems, retrieval-augmented generation, document analysis and multimodal workflows, but organizations using it for high-stakes factual tasks will still need retrieval, verification and human review. How a 276B model uses only 12B parameters at a time Inkling-Small is a sparse Mixture-of-Experts model. According to the model card published by Thinking Machines, its 42-layer decoder routes each token to six of 256 specialized experts, along with two shared experts that remain active for every token. That architecture helps explain the distinction between the model’s 276 billion total parameters and its 12 billion active parameters. The system retains a large pool of learned capacity but activates only a fraction of it during each inference step. It is also natively multimodal. Images, audio and text are projected into a shared representation and processed jointly by the decoder rather than being handled through completely separate external systems. Thinking Machines lists coding assistants, agentic applications, chatbots, RAG systems and other multimodal applications among its intended uses. The company also supports variable reasoning effort, allowing developers to increase or reduce the model’s test-time compute depending on the difficulty of the task. That gives engineering teams a direct way to balance quality, latency and cost across different workloads. Unfortunately, small does not mean it runs on a laptop Despite its name, Inkling-Small is not a consumer-scale model. The standard BF16 checkpoint requires at least 600 GB of aggregate GPU memory, according to Thinking Machines. The company lists two supported configurations: 4x NVIDIA B300 GPUs or 8x NVIDIA H200 GPUs. A quantized NVFP4 checkpoint lowers the requirement to roughly 180 GB of aggregate VRAM. Thinking Machines says that version can run in W4A4 mode on a single NVIDIA B300, or in W4A16 mode on two H200 GPUs. That rules out ordinary laptops, MacBooks, desktop gaming PCs and most developer workstations. Even heavily equipped local systems generally fall far short of the required memory. The practical deployment targets are enterprise GPU servers, cloud clusters and specialized inference providers. The “Small” label is therefore relative to Inkling, not to the broader universe of local models. Still, the reduction is meaningful. A model that approaches Inkling’s performance while needing substantially less aggregate memory can lower hosting costs, make capacity planning easier and widen the group of organizations capable of self-hosting it. For companies that want control over data, model behavior and fine-tuning, that smaller footprint may be more important than chasing the highest possible benchmark score. And of course, it being open source means that it will no doubt be rapidly quantized (made less precise but requiring less compute) and likely blended with other models to be made even smaller for consumer-grade hardware. Apache 2.0 is the gold standard for enterprise open source models The licensing may be as important as the benchmarks. Inkling-Small is released under Apache 2.0, one of the software industry’s most familiar permissive licenses. It generally allows organizations to use, modify, fine-tune, redistribute and commercialize the model, including inside proprietary products, provided they comply with the license’s notice and attribution requirements. That gives enterprises far more legal flexibility than many custom “open” AI licenses, which may include revenue thresholds, branding obligations, use restrictions or separate conditions for large-scale commercial deployment. The distinction is increasingly relevant as more AI companies publish model weights without using a conventional open-source license. Chinese AI darling Moonshot for example, made the weights of its frontier class Kimi K3 model available earlier this week under a custom "open" license that includes additional commercial conditions rather than the comparatively straightforward terms of Apache 2.0. For legal, procurement and platform teams, that difference can materially simplify adoption. Apache 2.0 does not eliminate the need to review acceptable-use policies, data provenance, regulatory exposure or downstream safety obligations. But it gives organizations a clearer starting point for building internal systems, shipping commercial products and maintaining modified versions of the model. A more repeatable model-development pipeline Inkling-Small also shows how quickly Thinking Machines has turned its first large model release into a repeatable engineering process. Thinking Machines researcher Horace He contrasted the two launches in a post on X : “Whereas I felt like it took a village to release Inkling, Inkling-Small felt much more routine 😆 We just took the pipeline used for Inkling, passed in a smaller model, and voila — new model! Inkling Small benefited quite a bit vs Inkling from some minor improvements, but there’s still so much more left in the tank...” The comment suggests the company is no longer treating each model as a one-off research project. Instead, it is building a reusable pipeline for pre-training, post-training, reinforcement learning, evaluation and release. Thinking Machines says Inkling-Small benefited from an improved pre-training data mix, changes to the machine-learning recipe and on-policy distillation using Inkling as a teacher. The team then continued agentic coding reinforcement learning for two weeks. Mira Murati emphasized the same point in her own post, describing Inkling-Small as comparable to Inkling at one quarter of the size and highlighting that the weights were open and fine-tunable on Tinker immediately. How enterprises and AI builders should think about Inkling Small The company is also distributing full BF16 and NVFP4 checkpoints and supporting deployment through SGLang, vLLM, TokenSpeed, Unsloth and Hugging Face tooling. That combination gives developers several deployment paths: use an API, fine-tune through Tinker, rely on a third-party inference provider, or operate the model on private infrastructure. Inkling-Small is not a model that most individuals will download and run locally. But for businesses deciding between a very large flagship and a more manageable open-weight system, it presents a compelling compromise: nearly the same measured intelligence, stronger results on several coding and reasoning tasks, lower token pricing, a smaller hardware footprint and a license that permits broad commercial development. The broader signal may be just as important. Thinking Machines is showing that Inkling was not a one-time release. The company is already compressing its model family, refining its training pipeline and moving toward a cadence in which open-weight multimodal systems can be produced, improved and deployed more routinely.

VentureBeat·July 31, 2026·7 min read
AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost
AIRelease

AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost

To quote an ancient Jedi Master "Begun, the AI price wars have!" OpenAI is sharply reducing the prices of two models in its GPT-5.6 frontier series, cutting GPT-5.6 Luna, the smallest and fastest model in the series, by 80% and GPT-5.6 Terra, the mid-tier model, by 20%, while adding a premium Fast mode for its flagship GPT-5.6 Sol model. The cuts place Luna much closer to the lowest-cost commercial models in the market and arrive just a few days after Anthropic released its highly performant Claude Opus 5 at the same price as Opus 4.8, and Google introduced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite , two rival models built around lower inference costs, faster execution and more efficient agent workloads. O penAI is successfully undercutting Google's price per intelligence and attempting to sway Anthropic users, who may not mind paying more, with a speed boost. OpenAI says Luna will now cost $0.20 per million input tokens and $1.20 per million output tokens, for a combined input-plus-output price of $1.40 per million tokens. Terra will cost $2 per million input tokens and $12 per million output tokens, for a combined price of $14. Pricing for Sol Standard remains unchanged at $5 per million input tokens and $30 per million output tokens. OpenAI is also adding Sol Fast mode at twice the Standard price: $10 per million input tokens and $60 per million output tokens. The company says Fast mode delivers up to 2.5 times the throughput without changing the model’s underlying intelligence. OpenAI co-founder and CEO Sam Altman took to X to announce the changes as "major price cuts today." VentureBeat Frontier AI model API pricing comparison Model Input ($/1M) Output ($/1M) Total ($/1M) Source MiMo-V2.5 Flash $0.10 $0.30 $0.40 Xiaomi deepseek-v4-flash $0.14 $0.28 $0.42 DeepSeek deepseek-v4-pro $0.435 $0.87 $1.305 DeepSeek GPT-5.6 Luna $0.20 $1.20 $1.40 OpenAI MiniMax-M3 $0.30 $1.20 $1.50 MiniMax LongCat-2.0 — limited-time promo $0.30 $1.20 $1.50 LongCat Gemini 3.1 Flash-Lite $0.25 $1.50 $1.75 Google Qwen3.7-Plus $0.40 $1.60 $2.00 Alibaba Cloud MiMo-V2.5 $0.40 $2.00 $2.40 Xiaomi Gemini 3.5 Flash-Lite $0.30 $2.50 $2.80 Google LongCat-2.0 — standard $0.75 $2.95 $3.70 LongCat MiMo-V2.5 Pro (≤256K) $1.00 $3.00 $4.00 Xiaomi GLM-5.2 $1.40 $4.40 $5.80 Z.ai Grok 4.5 $2.00 $6.00 $8.00 xAI MiMo-V2.5 Pro (>256K) $2.00 $6.00 $8.00 Xiaomi Gemini 3.6 Flash $1.50 $7.50 $9.00 Google Qwen3.7-Max $2.50 $7.50 $10.00 Alibaba Cloud Gemini 3.5 Flash $1.50 $9.00 $10.50 Google Gemini 3.1 Pro Preview (≤200K) $2.00 $12.00 $14.00 Google GPT-5.6 Terra $2.00 $12.00 $14.00 OpenAI GPT-5.4 $2.50 $15.00 $17.50 OpenAI Kimi K3 $3.00 $15.00 $18.00 Moonshot AI Gemini 3.1 Pro Preview (>200K) $4.00 $18.00 $22.00 Google Claude Opus 5 $5.00 $25.00 $30.00 Anthropic GPT-5.5 $5.00 $30.00 $35.00 OpenAI GPT-5.5 Instant (chat-latest) $5.00 $30.00 $35.00 OpenAI Sakana Fugu Ultra (≤272K) $5.00 $30.00 $35.00 Sakana AI GPT-5.6 Sol — Standard mode $5.00 $30.00 $35.00 OpenAI Claude Fable 5 / Claude Mythos 5 $10.00 $50.00 $60.00 Anthropic GPT-5.6 Sol — Fast mode $10.00 $60.00 $70.00 OpenAI Pricing is shown per one million tokens. Total cost is calculated as input price plus output price. Cached-input pricing is excluded to keep the comparison consistent across providers. OpenAI moves Luna into the low-cost tier The most consequential change is the Luna price cut. When OpenAI introduced the GPT-5.6 series, Luna was priced at $1 per million input tokens and $6 per million output tokens, for a combined total of $7. The new pricing reduces that combined figure to $1.40. That places Luna below Google’s Gemini 3.5 Flash-Lite, which costs a combined $2.80 per million input and output tokens, and far below Gemini 3.6 Flash at $9. Luna also now costs less than OpenAI’s own GPT-5.4 and Terra models by a wide margin. It is not the cheapest model in the broader market. Xiaomi’s MiMo-V2.5 Flash, DeepSeek’s flash model and several other APIs remain less expensive on a pure token basis. But the reduction brings an OpenAI frontier-series model into direct competition with the market’s low-cost inference tier. OpenAI says the GPT-5.6 series represents its frontier model family, with Sol positioned at the top of the lineup, Terra as the middle tier and Luna as the smallest and fastest option. The lineup was initially released in late June 2026 through a limited rollout by U.S. government request, before broader access, with each model intended to offer a different tradeoff among intelligence, latency and cost. Sol is aimed at the most complex reasoning-heavy and agentic workloads, including advanced coding, multi-step planning and tool-using systems, while Terra is designed for general production use where a balance of capability and efficiency is required. Luna is positioned for high-throughput, low-latency tasks such as summarization, classification, routing, and lightweight real-time assistants where cost per request is the primary constraint. Terra drops to match Google’s Gemini 3.1 Pro pricing Terra’s 20% reduction moves its combined price from $17.50 to $14 per million tokens. At that level, Terra now matches Google’s Gemini 3.1 Pro Preview pricing for context windows of 200,000 tokens or less. It also undercuts OpenAI’s GPT-5.4, which remains priced at $2.50 per million input tokens and $15 per million output tokens, offering the same intelligence for about 1/13th the cost, as Krea AI's Nic Dunz noted on X : The adjustment creates a wider separation between OpenAI’s three GPT-5.6 tiers. Luna costs one-tenth as much as Terra on a simple combined input-plus-output basis, while Terra costs 60% less than Sol Standard. Sol Fast moves in the opposite direction. At a combined $70 per million tokens, it is the most expensive model configuration in the comparison below, reflecting OpenAI’s decision to charge a premium for latency-sensitive workloads rather than lower Sol’s base price. Cuts follow Google’s low-cost Gemini releases and Anthropic's Claude Opus 5 OpenAI’s pricing changes come only about a week and a half after Google introduced its own low-cost Gemini 3.6 Flash and Gemini 3.5 Flash-Lite . Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens. Google framed both models around the economics of agent deployment, arguing that lower token usage, fewer reasoning steps and reduced tool calls could lower the total cost of long-running software engineering and knowledge-work tasks. Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching as high as 65% on some long-horizon engineering workloads. Gemini 3.5 Flash-Lite is positioned as the fastest model in Google’s 3.5 series. However, OpenAI's models are more performant than Google's, according to third party analysis outfits like Artificial Analysis , with even the Luna model outperforming Gemini 3.6 Flash and the older Gemini 3.1 Pro model, making the cost-per intelligence much more favorable to OpenAI. As AI coding startup Cognition noted on X , GPT-5.6 now "sits on the pareto curve of price/performance efficiency," posting an animation of the GPT-5.6 series moving left on a chart representing intelligence on the y axis and cost on the x, showing that the models now offer among the most superior intelligence for lowest cost on the market. And yet, rival Anthropic's Claude Opus 5 remains about as performant as GPT-5.6 Sol, yet is 6% cheaper. The model costs $5 per million input tokens and $25 per million output tokens—the same rates as Opus 4.8—but Anthropic says it delivers nearly all the intelligence of its more expensive Fable 5 model at roughly half the cost. Unlike OpenAI’s Luna and Terra changes, Anthropic did not reduce the Opus API sticker price. Instead, it effectively lowered the price per unit of capability by replacing Opus 4.8 with a more capable model at the same $30 combined input-and-output rate. Anthropic also added an adjustable effort setting that allows developers to trade reasoning depth for speed and token savings. That distinction matters for enterprise buyers. OpenAI is directly cutting per-token rates, Google is pairing lower prices with reductions in token use and tool calls, and Anthropic is emphasizing stronger task performance at an unchanged price. All three approaches target the same operational metric: the total cost of completing production work, rather than the advertised cost of an individual token alone. The timing highlights how quickly pricing has become a competitive lever among frontier model providers. OpenAI’s response does not introduce a new model generation. Instead, it changes the economics of deploying models that were released only recently. The market shifts from model access to model economics The cuts indicate that access to frontier-level capability is no longer the only point of competition. The next question for enterprises is how cheaply and predictably those models can run in production. OpenAI is still not the lowest-priced provider on a pure token basis. But Luna’s 80% reduction materially changes its position, moving it from the middle of the market into a pricing tier populated by smaller models from Google, Xiaomi, DeepSeek, MiniMax and other vendors. That matters most for high-volume applications, where relatively small differences in token pricing can compound across coding agents, document systems, internal search tools and automated workflows. OpenAI’s latest move therefore looks less like a routine adjustment and more like a repositioning of the GPT-5.6 series. Sol remains the premium option, Terra moves closer to competing pro-tier systems, and Luna becomes the company’s direct answer to the industry’s growing low-cost model segment.

VentureBeat·July 30, 2026·8 min read
Mastercard spent decades training its fraud system to see bots as thieves. Now bots are the ones doing the buying.
AISecurity Advisory

Mastercard spent decades training its fraud system to see bots as thieves. Now bots are the ones doing the buying.

Every time a Mastercard gets tapped, the network has less than a tenth of a second to judge how likely the purchase is to be fraudulent. It made that call across 175 billion transactions last year. Now the buyer on the other side of that judgment is starting to change, and Greg Ulrich, the company's chief AI and data officer, spelled out the consequence for the VB Transform 2026 audience in Menlo Park on July 14. "We've built a bunch of risk rules over time that were intended to stop a bot from transacting," Ulrich said. "Now we need to enable the bot to transact, so that requires a change to our risk framework and our risk rules." Ulrich joined Mastercard eleven years ago when an analytics company he worked at was acquired, and said trust struck him from day one on the job. "It's what enables a merchant that's never met you to accept payment and ensure that they're going to get paid. It's what enables you as a consumer to transact and ensure that things are going to work out in a trusted, secure way. And if something goes wrong, there's a safe and secure path for a dispute and to resolve this," he said. 175 billion transactions, scored in under 100 milliseconds He took the audience inside each of those calls. "When you tap your Mastercard to pay for a product or service, we're providing a score to that transaction," he said. "We have under 100 milliseconds to look at that and give a score from zero to 999 about how likely is that to be fraudulent or real. And we pass that on to the issuing bank." Generative AI widened what that score can see. "Because we have new technology, we can bring in more data, we can bring in more context, and now we're finding that we can identify 300, 400% more fraudulent transactions at those high-risk bands," Ulrich said, without adding friction or false positives for consumers. The company's Safety Net system has stopped more than 70 billion fraudulent transactions, he told the audience, and Mastercard is building its own transformer model on its transaction data as a foundation for new safety, security, and personalization solutions. VentureBeat's Beyond the Pilot podcast took that production fraud stack apart in detail earlier this year. A third of the services business already runs on AI The business stakes reach past fraud. About 40% of Mastercard's company is now based on services, Ulrich said, including marketing services; fraud, safety and security; and business intelligence. "A third of those are predicated on AI, and those are growing at a much faster clip than everything else," he said. One line he returned to all session went further. "What's going to enable AI to continue to scale is not the capabilities of the agents, it's how much we trust those agents to do on our behalf as a consumer, as a business, as a financial institution, or otherwise," he said. Five layers stand between agents and the network Agentic commerce changes the object being secured. "Instead of a single atomic transaction where I say go buy something, I'm effectively delegating authority, or a consumer's delegating authority, a business is delegating authority," Ulrich said. "And when that happens, it's a much more complicated transaction." Trust, in turn, has a precondition. "The only way it's going to work with trust is if we can identify what was the intent, what are the behaviors, what are the constraints that were intended in that transaction." Ulrich walked through five layers Mastercard has built against that problem. Identity comes first. "I want to make sure I can understand not just who the consumer is, but who the agent is, that I combine them together and that I have KYA or know your agent, that I'm validating that it's legitimate technology, that it's a legitimate agent," he said. "We can register it into our system." Verifiable intent settles the "wrong-Nikes" problem Verifiable intent is second, a tamper-proof cryptographic record of the original instructions that travels with the transaction. "If you've asked for Nike black Nikes in size 12, but you got them on a final sale and they're not returnable and that wasn't in your instruction, there's a way to look at that in an objective and clear way on the back end," he explained. Controls form the third layer, defining which merchants an agent can buy from, at what limit, and under what constraints. Execution runs through Mastercard Agent Pay , which carries "the tokenization, authentication, the acceptance framework embedded within it" and has launched with Microsoft, OpenAI, Google, and others, Ulrich said. Intelligence is the fifth layer, spanning risk rules, insight tokens that grant "consented or permissioned access to insights" for personalized recommendations, and monitoring through Recorded Future to identify threat actors in the system. The bigger prize is a procurement agent with a budget Consumer purchases are where agentic commerce started. Ulrich pointed the room past them, to business-to-business procurement as the larger opportunity. His example was a manufacturer that wants an always-on assembly line, with an agent that manages inventory levels, tracks when stock runs low, replenishes automatically, and understands the budget and the approved suppliers. "When you can start enabling that, you require those same five layers for that type of transaction," he said. Making it work across companies multiplies the parties that have to trust each other. "You need clear standards for identity, you need clear standards for intent, you need these to work across. You're gonna have a procurement agent, a supplier agent, a banking agent. They're all gonna need to communicate to enable this to happen in an autonomous way, and that's gonna require really scaled trust infrastructure." Powerful new models, same security motion Mastercard sat in the early wave of Project Glasswing with Anthropic's Mythos model, and worked with OpenAI's GPT-5.5-Cyber , he said. "What we've seen from both of those is incredibly powerful models finding new vulnerabilities in the ecosystem that were difficult to detect previously, but it's really a new tool as opposed to a new motion," Ulrich said. Inside the company, the chief security officer leads that work. A dedicated team has prioritized the most critical assets, runs them through the models routinely, tracks findings by high, medium, and low severity, and uses the same technology to handle patches. Ulrich said the approach has already been extended out, and that Mastercard is working to make the same architecture and patching available to others as well. What Mastercard would build differently after 14 months "The guardrails, the security, all this stuff has to be embedded at the front end. These can't be things that we're adding on at the back end. That's lesson one. Lesson two is you have to be operating for scale, and the other one is around observability and accountability matter as much as the intelligence," Ulrich said, counting off what building inside Mastercard taught the team. The company built what he described as an agentic factory, an operating system with the compliance, the observability, and the guardrails built in rather than bolted on per agent. Model drift, once tracked manually by dedicated teams, is now automated into that factory. Asked by an audience member about the gotchas, Ulrich did not soften the pilot-to-production trap. "If you're trying to extend that and then add guardrails in as you're extending it, once you've already built it, I think you're doomed to fail," he said. Mastercard built a series of agents last year for its 4,000 consultants, covering deep research, text to SQL, Excel, and PowerPoint, tools that by his account did not exist at the level Mastercard needed. Were the company starting today, Ulrich said, it would build them fundamentally differently. "I don't know that we anticipated when we built things fourteen months ago that we would be rethinking the fundamental architecture and the approach already." Agentic identity joins KYB and KYC The identity layer is where Ulrich expects the market to move next. Inside Agent Pay, Mastercard authenticates the consumer the way it does in traditional e-commerce and binds the agent to that person. "Outside of that framework, I think there will be open standards to identify who an agent is and bind the agent with the consumer," he said. "And then we can tie that with verifiable intent." VentureBeat's June 2026 Pulse research points at the same gap. Only 32% of the 107 qualified enterprise respondents give every agent its own scoped, managed identity , and just 12% include an agent-identity product in their consideration set. He called identity "one of the faster-growing ecosystems," noting Mastercard has been expanding there organically and inorganically for about six or seven years, with the work now spanning "agentic identity as well as the traditional KYB and KYC identity." The risk rules that keep bots off the network came out of more than two decades of applying AI to those transactions. The rewrite, for the agents Mastercard now wants to let in, is already underway on the same network that scored 175 billion of them last year.

VentureBeat·July 30, 2026·8 min read
Zero-Day
↘2K
💬

Top Discussion

HN

Hacker News

“GPT-5.5's API pricing is reshaping how startups build AI products”

14.1K932 comments
View discussion→

Filters

Time
Categories
Sources
Content Type