Warning: foreach() argument must be of type array|object, null given in /home/aymanweb/theinvestorsnews.com/wp-content/plugins/wp-simple-firewall/src/Controller/Database/DbCon.php on line 251
Nvidia Took Claude From 30% to 100% – Without Building a Better Model - The Investors News
Thursday, August 27, 2026
Daily News

Nvidia Took Claude From 30% to 100% – Without Building a Better Model

EA Builder

On one of AI’s toughest interactive tests, Claude solved fewer than one in three challenges.

Then Nvidia (NVDA) surrounded the same model family with better memory, outside tools, a structured work loop, and a supervisor that stepped in whenever Claude got stuck.

The system solved every challenge – and its score jumped from 30.2% to 100%.

Nvidia changed other parts of the test setup, too, so this was not a perfect apples-to-apples comparison. Still, the result points toward a major shift in the AI market.

Today’s best models may already contain more useful intelligence than their surrounding systems allow them to express.

The companies that turn that capability into dependable, cost-effective workflows could become some of the next major AI winners.

Why AI Agents Still Struggle With Long-Horizon Tasks

Today’s AI agents are pretty good at sprints. Marathons are where they lose the plot. 

Give Claude or GPT a short assignment. It can browse the web, write some code, pull information from another app, and come back with an answer.

But long projects are a different story. Imagine asking an agent to migrate a company’s financial software, redesign a supply chain, or optimize thousands of lines of GPU code. The work may require hundreds of decisions made over hours or days.

Frontier models still struggle with that kind of sustained work. They lose context. Repeat old mistakes. Chase dead ends. Occasionally become so confused that they damage the project they were supposed to complete – and we’re left to try to clean up the mess.

Nvidia’s Agentic Variation Operators system, or AVO, was designed to keep that from happening.

AVO gives the model persistent memory and its own repeated cycle of planning, acting, testing, and revising. Nvidia also added a supervising agent that watches the work and nudges the main agent when it gets stuck or starts exploring an unproductive path.

Think of it as a talented employee with a good project-management system and an experienced boss nearby.

The intelligence was already there.

Nvidia helped it stay organized long enough to finish the job.

AI Agent Infrastructure Is Becoming the Next Battleground

For the first few years of this boom, the model leaderboard commanded nearly all the attention.

Which lab had the best reasoning score? Which model had the largest context window or wrote the cleanest code?

Those are still important questions; better models still retain a huge competitive advantage.

But Nvidia’s experiment suggests that raw capability alone will not determine every winner…

A model’s surrounding system – its ‘harness’ – decides how the AI receives information, what tools it can use, how it remembers prior work, when it asks for help, and what happens after it makes a mistake.

That opens a much broader market.

Companies will need software that coordinates agents across long workflows. Memory systems that preserve the right information without dragging old tokens into every new task. Tools that route jobs toward the model best suited to handle them.

Security becomes more important, too – because the more freedom an agent receives, the more carefully companies have to control what it can access, what actions it can take, and when a human needs to get involved.

Then there is monitoring.

Businesses will want to know why an agent made a decision, how much the task cost, which models and tools it used, and whether its work can be trusted.

The model provides the intelligence. An increasingly valuable group of companies will help turn that intelligence into reliable work.

Better AI Agent Orchestration Can Make Smaller Models More Capable

Nvidia’s result is not the only evidence pointing in this direction.

Inherent Labs, a startup founded by former Google DeepMind researchers, recently unveiled an AI research agent called Faraday.

Faraday runs on a comparatively small 27-billion-parameter model. Despite that smaller engine, Inherent says Faraday beat much larger systems from Anthropic and OpenAI at one demanding research job: reproducing the results of published scientific papers. 

That does not prove smaller models are generally better than frontier systems. But it does show how much useful performance can come from the way an agent is trained, equipped, and organized.

Faraday does not try to build every capability from scratch. For coding work, it can call OpenAI’s Codex, much like a human researcher uses specialized software rather than reinventing every tool personally.

That is probably closer to how enterprise AI will develop.

Companies may use one model for high-stakes reasoning, another for routine document work, and a specialized coding system for software tasks. An orchestration layer will decide which tool gets the job and keep the full process moving.

The best enterprise agent may not have the smartest model at its center.

It may have the best team of models and the strongest system for managing them.

The Same Model Can Produce Very Different Economics

Performance is only half the equation, though. Companies also care what it costs to finish the job.

As Databricks CEO Ali Ghodsi explained to TechCrunch, two agent systems built around the same model can produce dramatically different bills. Choose the wrong setup, and the same task may cost roughly twice as much to complete.

A clumsy workflow sends a routine job to an expensive frontier model when a smaller one would do. Poor memory leads the system to reread huge amounts of old information again and again.

A better setup preserves the useful information, sends each task to the right model, and cuts out unnecessary loops.

That creates opportunities for the companies building agent software, memory systems, model-routing tools, security, monitoring, and developer platforms.

It also keeps the physical infrastructure underneath them busy.

In another Nvidia test, AVO tried more than 500 ways to improve a piece of GPU software, saved 40 different versions, and eventually beat a leading implementation by as much as 10.5%.

That result took far more than one prompt. The agent kept testing, checking, and revising until it found a better answer.

Every loop consumed more compute.

That is the central trade-off in agentic AI. Better systems can lower the cost of completing each job while making far more jobs worth automating.

Companies spend less per completed task. Total AI usage keeps climbing.

And as that market expands, more of the value begins flowing into the software and infrastructure that make the full system dependable.

What the Agentic AI Shift Means for Investors

The model race remains fierce. 

Cloud providers are building their own agent platforms. Startups are attacking memory, routing, monitoring, and security. Existing software leaders are adding agent tools to products already embedded inside large companies.

The market is beginning to form around the systems that make models productive.

So we’ve honed a portfolio for the market in front of us.

After sifting through more than 200 different AI-related recommendations, Louis Navellier, Eric Fry, and I have narrowed that large universe to roughly 20 stocks we believe deserve capital now, including a recommended position size for every holding.

Check out the newly rebuilt AI Revolution Portfolio nowbefore the page goes offline at midnight.

The post Nvidia Took Claude From 30% to 100% – Without Building a Better Model appeared first on InvestorPlace.

Source link

Share with your friends!

Leave a Reply

Your email address will not be published. Required fields are marked *

Solverwp- WordPress Theme and Plugin

Get The Latest Investing Tips
Straight to your inbox

Subscribe to our mailing list and get interesting stuff and updates to your email inbox.

Thank you for subscribing.

Something went wrong.