“AGI has arrived.”

That was Nvidia founder and CEO Jensen Huang’s verdict on OpenAI’s GPT-6 Astra. Artificial general intelligence broadly describes AI that matches or exceeds human capabilities across many tasks—but whether Astra qualifies remains disputed.

For business leaders, a more practical question matters:

What valuable work can we now delegate to AI—and how do we know it was done properly?

Influential voices see a shift

OpenAI president Greg Brockman describes the beginning of an “AGI era,” suggesting that people may eventually identify this period, and possibly Astra, as the turning point.

Aaron Levie, CEO of Box, offered an enterprise perspective. After early testing, he called Astra “the best model we’ve ever tested” on Box’s expanded, hardest evaluation set, highlighting its capabilities in complex knowledge work.

Microsoft’s Satya Nadella highlighted early customers already using Astra on Azure. That is evidence of deployment interest—not a declaration from Nadella that AGI has been achieved.

Not everyone agrees with Huang. AI researcher Gary Marcus challenged the announcement for lacking a sufficiently clear definition and supporting evidence. Prominent endorsements do not settle a scientific question.

Businesses do not need agreement about AGI before testing useful capabilities. They do need evidence before depending on them.

Box reported 77 percent overall accuracy for GPT-6 Astra versus 74 percent for GPT-5.6 Sol. On one NDA-review task, Astra scored 93 percent versus 69 percent. These are different evaluation scopes, not Shofield customer outcomes.
Box’s own test results, visualised by Shofield AI. The overall evaluation and individual NDA-review task have different scopes. Shofield AI chart; data: Box Blog

The relevant evidence is about real work

Box tested Astra on complex, multi-document business tasks. It reported 77% overall accuracy, compared with 74% for GPT-5.6 Sol. On one specific task—reviewing a non-disclosure agreement against company policy—the score increased from 69% to 93%.

These are Box’s results under its own test conditions, not guaranteed outcomes for Shofield customers or evidence that professional review is unnecessary.

The example matters because reaching a plausible conclusion was insufficient: the model also needed to identify the policy provision supporting it. That is closer to the requirements of professional work than simply producing convincing text.

OpenAI similarly emphasises Astra’s ability to execute multi-step workflows and produce usable documents, spreadsheets and presentations.

The business test should therefore be straightforward: does the system complete our workflow correctly, with less total effort and acceptable cost?

What this means when Astra is connected to Shofield

A powerful model supplies intelligence. An operational AI workforce also needs company knowledge, authorised system access, instructions, approval boundaries and completion checks.

Shofield brings together Company Brain, AI Workforce and Mission Command, alongside engines for sales, healthcare and cybersecurity, supported by implementation and managed operations.

For organisations using Astra-enabled Shofield workflows, the objective is to make advanced reasoning useful inside the business—not leave it as another disconnected tool.

The following are potential deployment examples, subject to enabled modules, authorised integrations and agreed scope.

Professional services: prepare work for review. A workflow could assemble client information, identify missing documents, flag inconsistencies and prepare structured summaries. The goal is less coordination and preparation work, leaving qualified professionals to focus on judgement and client advice. Measure preparation time, completeness and rework—not document volume.

Sales: give every opportunity a relevant next step. Evaluate Astra-enabled research, qualification and follow-up against better conversations and stronger progression towards purchase. The objective is not simply more outbound activity. Response time, follow-up completion and conversion to paid should matter more than message volume.

Healthcare operations: reduce administrative friction. Start with referral preparation, missing-information checks, draft communications and routing exceptions to named staff. These are administrative workflows, not permission for a general-purpose model to make clinical decisions. Measure processing time, backlog and manual interventions.

Cybersecurity: move from findings towards verified responses. A bounded workflow could interpret authorised evidence, organise findings and prepare remediation for approval. OpenAI identifies secure code review and patching as supported defensive uses, while more sensitive capabilities remain subject to additional safeguards and access arrangements.

Why managed services matter

A stronger model does not automatically establish responsibility for integrations, failures, costs or changing behaviour.

Shofield’s Managed AI Operations proposition covers monitoring available performance, quality and cost signals; agreed support and incident responsibilities; evaluating changes before promotion; and expanding scope when a workflow demonstrates value.

For customers, that means distinguishing access to intelligence from accountability for how it performs.

A managed deployment should have an owner, acceptance criteria and an escalation process. Model updates should earn their place through testing. Greater autonomy should follow demonstrated reliability.

The economic objective is equally clear: reduce the cost of correctly completed work—not merely the cost of generating an answer.

For example, saving ten minutes across 600 monthly cases would release 100 staff hours. That is illustrative, not a Shofield result or automatic cash saving. Review, usage and operating costs still need to be included.

An AI-proposed action passes identity, permission and budget checks. Actions requiring human approval wait for a decision before authorised execution. Requests, decisions and results are recorded outside the model, and an operating recovery path can pause, contain or restore work.
A conceptual operating architecture: authority is enforced by the surrounding system. This is Shofield’s recommendation, not a screenshot or benchmark claim. Shofield AI

Greater capability requires stronger control

OpenAI’s system card reports reduced visibility into Astra’s chain-of-thought reasoning, alongside improved results in respecting safety and security restrictions. Both findings matter.

Our operational recommendation is to enforce permissions, spending limits and approvals through the surrounding system—not rely solely on the model to police itself. Consequential actions should leave evidence, and sensitive changes should have a recovery path.

More capable AI should earn more responsibility through evidence, not receive unlimited authority because of a headline.

Start with one valuable workflow

The AGI debate will continue. Your organisation can focus on execution.

Choose one consequential but bounded workflow. Establish the baseline. Connect the necessary information and systems. Define approval boundaries. Measure the result before expanding.

Contact us to identify a high-value workflow and assess an Astra-powered Shofield deployment, supported by implementation and managed AI operations.

Leading AI. Delivered.

SOURCES & FURTHER READING