Signal65’s PINNACLE benchmark gives technology leaders a new way to apply the production gates in Digital CxO’s No Maintenance Window report: measure completed work, test it against the data the enterprise actually has and price the result that comes out right.

Technology leaders have spent the last several years being told that choosing an AI model begins with a leaderboard. Find the model with the highest score, compare its token price with the alternatives and decide whether the gain in capability justifies the premium. That approach is understandable. It is also increasingly disconnected from the decision an enterprise must make.

A general-purpose leaderboard can tell a leader how a model performed on a standardized test. It cannot tell a CIO whether an agent will complete a specific enterprise workflow, survive contradictory company data, operate at the required volume and deliver an accepted result at a defensible cost. Those are not secondary implementation questions. They are the difference between an impressive demonstration and durable production.

That distinction sits at the center of our new Digital CxO special report, No Maintenance Window: What Technology Leaders Should Keep, Change and Wait On as Agentic AI Moves Into Production. The report argues that the enterprise objective is not maximum autonomy or maximum AI adoption. It is completed, accepted work that can be repeated, observed, secured, recovered, economically defended and owned.

Signal65’s launch of PINNACLE gives that leadership argument a new set of measurements. PINNACLE does not ask only whether a model appears intelligent. It runs multi-step enterprise jobs against filesystems, databases and Python environments, grades the completed work in code against an answer key held outside the agent’s sandbox, measures the capacity of the infrastructure underneath it and calculates what each correct task costs.

The Wrong Unit of Measurement

The need for a different instrument is visible in the buying data. ETR research cited in the PINNACLE launch analysis found that 67% of enterprises building AI identified data quality as their largest challenge, 56% were already running multiple models in production and only 7% saw model families as highly differentiated. Enterprises are assembling multi-model estates while struggling with the data those models must use, yet the market continues to present model selection as a contest with one universal winner.

PINNACLE changes the unit from an abstract model score to correct work. Every job is evaluated as a complete deliverable rather than a collection of promising intermediate steps. Every hosted model is priced from the same runs used to grade the work, so failed attempts remain part of the bill. Open models are measured on actual accelerator nodes, allowing buyers to compare the economics of renting API intelligence with running weights on infrastructure they control.

That directly reinforces the first gate in No Maintenance Window: define the completed result that has value before selecting the technology. If the organization cannot state what an accepted outcome looks like, a higher benchmark score gives it no reliable basis for deciding whether the workflow succeeded.

The 28-Point Reality Tax

The most important PINNACLE finding is not which model occupies the top row. It is what happens when the same model meets the data an enterprise actually has.

The launch results cover 44 configurations from 30 base models. Each configuration ran the same jobs under two conditions. The first used organized and governed data. The second used data with the duplicates, partial migrations and contradictions that accumulate inside a working company. Forty-three of the 44 configurations lost ground in the second condition. The median configuration surrendered about 28 percentage points of whole-workflow completion, and the largest decline reached 64 points. Nine configurations completed at least 95% of the jobs on governed data; only two cleared 95% on the as-found version.

It is tempting to treat that as another model result. It is also a leadership result. Data readiness is not an infrastructure chore to be completed after a model has been selected. It is part of the production system, and its effect varies by model. A company may be able to engineer around a large data-sensitivity gap by governing and cleaning the relevant sources. It cannot engineer around a model that performs poorly in both conditions.

No Maintenance Window makes production-grade data and context its third durable-production gate and evaluation and acceptance its fifth. PINNACLE shows why those gates belong together. Testing a model against a polished benchmark corpus can demonstrate capability. Testing the completed workflow against the conditions in which it will operate reveals readiness.

Completed Work Is the Only Honest Denominator

PINNACLE’s second contribution is economic. The AI market typically advertises input and output token prices, but an agent does not read a prompt once and produce one answer. It retrieves material, calls tools, revises state and rereads a growing context across dozens of rounds. In Signal65’s testing, input represented 65% to 91% of the hosted-model invoice. Pricing the work from output tokens alone understated the cost of a correct task by three to 11 times.

A correct enterprise task using GPT-5.6 Sol cost $1.21 in the measured runs, not the dime suggested by simple sticker-price arithmetic. The best midsize open models running on an eight-GPU B300 node delivered a correct task for roughly eight to 15 cents and remained less expensive than the frontier APIs down to approximately one-quarter node utilization in Signal65’s model.

That does not make self-hosting automatically cheaper. A complete enterprise calculation also includes engineering, platform services, data movement, security, monitoring, human review, exceptions, integration maintenance and recovery. PINNACLE captures more of the bill than token-price comparisons do, but it is still an input to a fully loaded business case rather than the entire business case.

The seventh gate in No Maintenance Window is completed-work economics. Leaders should compare the total cost of the human, agentic and deterministic workflow with the value of accepted outcomes. Calls attempted, tokens consumed and documents generated may help explain the bill. They are not the denominator that determines whether the investment worked.

There Is No Best Model Without a Job

The launch rankings themselves reinforce the same point. Claude Opus 5 led PINNACLE’s overall table, but GPT-5.6 Sol led its IT Professional persona. GLM-5.2, an open-weight model, led Customer Operations because that persona placed greater weight on refusing to invent an answer that the underlying documents did not contain. The leading closed models fabricated unavailable answers more frequently than several of the strongest open configurations in this particular test.

Those results are not blanket recommendations. They show why the model decision must follow the workflow and its failure costs. A customer-service agent that confidently invents an unavailable answer presents a different risk from an internal research assistant that escalates uncertainty. A model that maximizes overall completion may not be the model that best satisfies the acceptance threshold for a particular role.

The same boundary applies to the finding that Chinese open-weight models occupied six of the launch table’s top 10 positions. That is a meaningful result for the enterprise agentic work PINNACLE currently measures. It is not a declaration about general intelligence, national AI leadership or every enterprise workload. Security, provenance, regulation and supply-chain requirements remain separate parts of the procurement decision.

A Better Benchmark Is Still Not the Enterprise

PINNACLE should not become a new oracle that simply replaces yesterday’s leaderboard. Its launch suite uses synthetic environments so that the answer key can be known and the tests cannot be memorized. It measures 280 agentic samples per configuration, with resolution stated at roughly five points for an environment score. The harness is not currently open for anyone to reproduce independently, and the first hardware-capacity results compare generations within a vendor rather than NVIDIA directly with AMD.

Those limitations matter, but so does the discipline of stating them. PINNACLE does not claim to predict every on-the-job outcome, cover whole occupations or measure general intelligence. The responsible enterprise response is therefore not to adopt its overall ranking as a procurement list. It is to adopt its measurement logic and apply it to the company’s own workflows, data conditions, failure thresholds and economics.

Keep, Change and Wait Now Has an Instrument

No Maintenance Window asks technology leaders to make three decisions. PINNACLE helps make each one more concrete.

Keep the disciplines that make production dependable: reliability engineering, least privilege, data lineage, modular architecture, change management, incident response and named human accountability. The messy-data and fabrication findings show what happens when context quality and acceptance criteria are treated as assumptions.

Change how success is measured. Move beyond leaderboard rank, token prices, prompts completed and hours ostensibly saved. Establish a completed-work scorecard that includes whole-workflow accuracy, exception rates, human review, recoverability and cost per accepted outcome. Retest when the model, prompt, tool, data or permissions change.

Wait before granting authority that the organization cannot observe, constrain or recover. A model should earn broader autonomy by demonstrating stable results under realistic data conditions, bounded behavior, acceptable economics and tested recovery. Waiting in this sense is not delay. It is staged authority based on evidence.

PINNACLE will not make those decisions for a CIO, CTO, CDO, CISO or CAIO. It does make the inadequacy of the old questions harder to ignore. The next AI procurement conversation should not begin with, ‘Which model is number one?’ It should begin with, ‘Number one at what work, on whose data, under what controls and at what cost per accepted result?’

There is still no maintenance window for the agentic transition. Leaders now have one more reason to insist on a production gate before allowing it to accelerate.

Read the full report

Download DigitalCxO’s special report, No Maintenance Window: What Technology Leaders Should Keep, Change and Wait On as Agentic AI Moves Into Production.

Explore the Signal65 PINNACLE benchmark and methodology at pinnacle.signal65.com.

Disclosure: DigitalCxO and Signal65 are both part of Futurum Group. Signal65 developed PINNACLE in partnership with Kamiwaza. NVIDIA and AMD provided hardware, engineering input and support under Signal65’s published governance rules; vendors have no approval or veto rights over published results.