Methodology · version 2026-09-04

Agentic Readiness: how we score what an AI agent can actually do

This is the complete rubric behind the Agentic Readiness score (0–100) used in our LLC formation comparison. It is vendor-neutral: any company — including ours — can be scored with it, and any company can reproduce its own score from public evidence.

Version 2026-09-049 criteria · 100 points19 providers scoredEvidence-only

Corpus authored this rubric and is scored by it. We score 19 providers including ourselves, and we publish every provider’s per-criterion breakdown below — ours included, with the criteria we lose points on.

If you believe your score is wrong, send us the public URL that shows it. We will re-verify against that source and correct the dataset. We have no affiliate relationship with any company scored here.

Why “has an AI chatbot” is not “agentic”

A chatbot answers questions. An agentic system lets software — an AI agent, an employee’s automation, a partner’s app — do the work: discover what is possible, learn the requirements, act, and observe the result.

A formation service is agentic to the degree that an external AI agent can complete this loop:

  1. Discover capabilities — a machine-readable description of what the service offers (an MCP server listing, an API spec, an llms.txt, an agents manifest).
  2. Retrieve requirements — structured, current data: what a given state requires, which fields, which fees.
  3. Collect structured inputs — intake captured as data, not free text.
  4. Validate them — machine-checkable validation before anything is filed.
  5. Initiate a workflow — create the order or engagement programmatically.
  6. Request human approval where designed — a gate the founder controls, not a hidden wall.
  7. Submit — the filing actually enters the state's process.
  8. Observe workflow state — status is queryable, by endpoint, webhook, or equivalent.
  9. Resume later — workflow state survives interruption.
  10. Retrieve evidence — confirmations, receipts, filed documents, an audit trail.

Most “AI-powered” formation sites stop at steps 1–2 with a support chatbot. That is AI-assisted, not agentic. The score exists to make the difference measurable rather than rhetorical.

The 100-point rubric

Criterion D — whether software outside the company can cause a filing to be submitted — carries the most weight alongside B, because it is the one that separates a description of a service from a usable one.

The nine Agentic Readiness criteria and their weights.
CriterionMaxWhat earns points
ANative conversational / agent interface150 none · partial = informational chatbot · high = an agent interface that can meaningfully perform the formation workflow, not just chat about it.
BMCP or equivalent agent protocol20A public MCP server or equivalent discoverable agent-tool protocol; more points for self-serve access, real tool coverage, and documented schemas.
CFormation API15A public or documented partner API with formation-relevant endpoints; more for public docs, sandboxes, and order-creating capability.
DCan an agent actually submit formation?15The decisive test: can software outside the company cause a filing to be submitted? Verified execution earns full points; internal-only AI earns partial; no path earns 0.
EProgrammatic status + resumable workflows10Status endpoints, polling, webhooks or events, and resumable session state.
FMachine-readable requirements and data10Structured access to state and filing requirements, fees, required fields, and schemas.
GAgent-compatible auth and payment5Can an external agent realistically complete the flow without an impossible human-only barrier? Systems are NOT penalised for intentional human approval gates designed for safety.
HHuman escalation / review5Human review, approval checkpoints, specialist escalation, support handoff.
IAuditability / provenance5Action history, state transitions, receipts, filing evidence, retrievable documents.

Scoring rules

  • Evidence onlyPoints require public evidence a reviewer could check: a URL, a doc, a verifiable endpoint. Internal capabilities that are not exposed publicly earn nothing — a company should get credit for what agents can actually use, not for what exists internally.
  • No unsupported negativesWe score YES / NO / PARTIAL / UNKNOWN. “We found no public evidence of X as of 2026-09-04” is the strongest negative claim we make. Absence of evidence is not evidence of absence, so unfindable features score UNKNOWN, not 0 — unless the company confirms the feature does not exist.
  • Date-stampedCapability and pricing data were verified on 2026-09-04. Scores drift; this page states its as-of date instead of implying it is live.
  • Scope penalties are explicitIf an agent interface covers only some states or entity types — a Wyoming-only MCP alongside a 50-state API, say — the limit is stated and reflected in the affected criterion, not hidden, and not double-counted in a second criterion.
  • Safety gates are not penalisedAn intentional human approval step does not cost points under criterion G. It cannot earn submission points that do not exist either: the rubric measures what an agent can cause to happen, and is deliberately silent on whether a company should let it.
  • ReproducibleEvery scored claim traces to a cited source in the evidence ledger carried in the dataset. Disagree with a cell and you can check it against the same URL we did.

Maturity levels

The level is a plain-language band for the score — useful because a two-point difference near the top of the rubric is not a meaningful distinction, while the gap between L1 and L4 is.

  • L0 · TraditionalNo meaningful AI or agent capability.
  • L1 · AI AssistedAI informs users; nothing executes.
  • L2 · Conversational WorkflowAI advances parts of the workflow inside the vendor's own app.
  • L3 · Agent-AccessibleExternal software can interact through APIs, MCP, or structured interfaces — fully open, or partner-gated.
  • L4 · Agent-NativeAgents can discover capabilities, initiate and continue workflows, submit formation, observe status, and operate with structured state and human approval.

Every provider’s score, criterion by criterion

This is the arithmetic. Each row sums to the published total, so you can check both the sum and any individual judgement against that company’s evidence.

Corpus is in this table under the same rules as everyone else, and its two lowest cells (D, whether an agent can submit, and E, status reporting) are the gaps it is trying to close.

Agentic Readiness by criterion, as of 2026-09-04.
CompanyA /15B /20C /15D /15E /10F /10G /5H /5I /5TotalLevel
Corpus1420101071055586L4
Doola91813159944485L4
Northwest Registered Agent1117368545463L3
Jupid1313383323250L3
LegalZoom919302815148L3
Tailor Brands7101075123348L3 (gated)
Swyft Filings001077614338L3 (gated)
Harbor Compliance40733215328L1
MyCompanyWorks00843105324L3 (gated; consumer side L0)
ZenBusiness90051003220L2
Rocket Lawyer60301014217L1
Stripe Atlas10300213313L1
Formations30000024211L1
Bizee (formerly Incfile)00000124310L0
BetterLegal40000103210L1
Clerky00300013310L0
Firstbase2010001329L1
LLC Attorney2000001328L0
Inc Authority0000001427L0

Machine-readable: the same numbers are in the criteria object of every company in /api/v1/compare/llc-services, keyless.

Known limitations of this rubric

A methodology that only lists its strengths is marketing. These are the weaknesses we know about:

  • It measures reachabilityThe rubric scores what is publicly reachable, which under-scores real capability sitting behind a partner agreement. We note gating explicitly rather than pretending the capability is absent, but a company with a strong private API will still score below one with a weaker public one.
  • Checkout behaviour is out of scopePreselected upsells and trial conversions are scored under pricing transparency, not here. A provider can be highly agentic and still have a checkout that surprises you.
  • Judgement bands are not perfectly calibratedCriterion G in particular produced adjacent scores on similar evidence across two providers in the 2026-09-04 pass. We kept the scores as originally judged and recorded the inconsistency rather than quietly restating one of them.
  • It is ours, and it will changeWe wrote the rubric, which is a conflict of interest no disclosure fully removes. Substantive changes will be versioned and dated on this page rather than applied silently to past scores.

Using this rubric

You may score yourself or anyone else with it, publish the result, and disagree with ours. If you do, cite the version date — a score without one is unfalsifiable, which is the failure mode this whole page exists to avoid.

The full comparison the scores feed is at /compare/llc-services. The dataset, including the per-criterion scores and every source URL, is at /api/v1/compare/llc-services with no key and no account. Corrections: support@corpuslaw.us.