Compliance

Mandatory AI Audits Arrive. The Audit Layer Is Already Broken

ai audit mandateillinois ai safety measures actcompliance evidence integrity
Hundreds of identical stamped seals lined up in rows in front of a bank of empty open filing drawers

On 6 July 2026 Illinois became the first US state to make independent third-party audits of frontier AI models a legal requirement. Four months earlier, the best-funded startup in the business of automating compliance audits was accused of producing 493 near-identical SOC 2 reports and was removed from Y Combinator. Those two facts belong in the same sentence, and almost nobody is putting them there.

The mandate is real and the direction is right. The problem is arithmetic. A legal duty to be audited creates demand for audit capacity that does not exist, demand gets met by automation, and we have just watched, in public, what automated compliance evidence looks like when nobody is checking the evidence. If your job is producing audit trails rather than consuming them, this is the year to make sure yours would survive being read.

What Illinois actually passed

SB 0315, the Artificial Intelligence Safety Measures Act, was signed by Governor JB Pritzker on 6 July 2026. It takes effect on 1 January 2027, with the audit obligations following a year later on 1 January 2028. Covered developers are those that earned more than $500 million in revenue in the preceding year, which is a threshold narrow enough that the list of affected companies is short and well known.

Two duties matter. Covered developers must implement, follow, and publish a frontier AI framework, meaning documented protocols for managing catastrophic risk. And they must retain an independent third party to perform an annual compliance audit against it. Enforcement sits exclusively with the Illinois Attorney General, with civil penalties up to $1 million for a first violation and up to $3 million for subsequent ones, per Skadden and DLA Piper.

The audit requirement is what separates Illinois from its neighbours in the same policy wave. California's Transparency in Frontier Artificial Intelligence Act and New York's RAISE Act both aim at the same class of developer and both stop at disclosure. Illinois is the first to say that publishing a framework is not enough and somebody independent has to check.

Three state AI laws, one difference that matters

LawCore dutyIndependent auditEnforcement
Illinois AISMA (SB 0315)Publish and follow a frontier AI frameworkYes, annual, from 1 Jan 2028Illinois AG, up to $1M / $3M
California TFAIATransparency and incident reportingNoState enforcement, disclosure-based
New York RAISE ActSafety protocol publicationNoState enforcement, disclosure-based

Read the right-hand column as a forecast rather than a scoreboard. Disclosure regimes tend to acquire verification once the first embarrassing disclosure turns out to have been wrong. Illinois has simply arrived there first, and the compliance industry now has eighteen months to build the capacity to answer it.

The part nobody wants to connect: the audit layer already failed

On 22 March 2026 an anonymous account published a detailed allegation against Delve, a Y Combinator company that had raised $32 million at a $300 million valuation selling automated SOC 2 compliance. The central claim was that the automated audits were not being audited. Of 494 SOC 2 reports examined, 493 were said to be near-identical, carrying the same paragraphs, the same grammatical errors, and the same nonsensical control descriptions, with only the company name and logo changed.

The follow-on allegations were worse: fabricated evidence of controls, auditor conclusions generated before any auditor had looked at client data, and unaccredited certification mills rubber-stamping the output. Around 3 April 2026 Y Combinator removed Delve from its directory and asked the founders to leave the programme. Delve's own position is that it does not issue compliance reports at all, describing itself as an automation platform that gives independent accredited auditors access to information, with customers free to choose their own.

Take the company's defence entirely at face value and the structural point survives intact. A platform that generates the evidence, formats the report, and routes the work to auditors it recruited is sitting on every side of a transaction whose entire value comes from independence. That is not a claim about anyone's honesty. It is a description of an incentive, and incentives are what controls exist to constrain.

Why this lands on firewall and network teams first

Frontier AI developers are not the audience for most of this site, and the Illinois threshold excludes almost everyone reading it. The relevance is the pattern, and the pattern is one that firewall change management has been living with for a decade before AI made it fashionable.

An audit report is a claim. The audit trail is the thing. When the two diverge, the report is what gets filed and the trail is what gets subpoenaed. Every organisation that has ever assembled an ISO 27001 or PCI evidence pack the week before an assessment knows the difference between a control that operates and a control that can be shown to have operated, and knows which of the two survives a serious auditor.

The failure mode is familiar. A spreadsheet of firewall rules is a report about the estate, not a record of it, which is exactly why spreadsheet-based firewall tracking fails audits the moment anyone asks who approved a specific change on a specific date. The equivalent question for an AI compliance report is who verified this control, on what evidence, and when, and 493 identical documents cannot answer it for 493 different companies.

What actually survives an audit that is trying

The distinction worth internalising is between generated evidence and captured evidence. Generated evidence is produced at report time by a system that is describing what it believes to be true. Captured evidence is recorded at the moment the thing happened, by the system that did it, with the identity of whoever authorised it attached.

  • Provenance beats presentation. A record that names the change, the requester, the approver, and the timestamp is worth more than any number of well-formatted control narratives.
  • Detect divergence, do not assert its absence. A running estate drifts from its documented state, which is why policy drift detection is evidence and a signed statement of compliance is not.
  • Recertification has to be an event with a date. If nobody can say when a rule was last confirmed as still needed, the recertification cycle exists on paper only.
  • Uniqueness is a signal. Two organisations with genuinely different estates should not produce evidence packs that differ only in the logo. If yours could be swapped with a peer's, it is describing a template rather than your network.

None of that is new advice. It is the same discipline behind an ISO 27001 firewall audit and behind the NIS2 evidence expectations for 2026. What is new is the volume of compliance work about to be created by statute, and the certainty that most of it will be met by tooling rather than by people.

The forecast worth planning against

Between the Illinois mandate, the EU's own timetable, and the disclosure laws already live in two other states, the demand for auditable AI documentation over the next twenty-four months will exceed the supply of qualified people willing to sign their name to it by a wide margin. Software fills that gap. Some of that software will be excellent, and some of it will be a template engine with a compliance logo, and the buyer will find it very hard to tell the difference from the outside.

The practical response is unglamorous. Ask any compliance vendor whether it generates the evidence, reviews the evidence, or both, and treat both as a finding rather than a feature. Ask to see two reports it produced for different clients. Keep your own captured records independent of whatever platform assembles them, because a vendor's collapse should not take your audit history with it. The same reasoning applies to mapping AI Act obligations onto real security controls, where the mapping is only as credible as the underlying evidence it points at.

The uncomfortable summary

Illinois did the right thing and did it first. Requiring an independent look at systems that their own builders describe in terms of catastrophic risk is not an overreach. But a mandate is a demand signal, and the market that will answer it has just shown, at scale and in public, that it is capable of producing compliance that means nothing.

Anyone whose organisation will be audited on anything in the next two years should read the Delve allegations as a preview rather than as gossip. The question an auditor should be asking, and the one you should be asking yourself first, has not changed since long before AI entered the room: not whether you have a report, but whether you can produce the record that report was supposed to be based on. If that record lives in a spreadsheet or in a vendor's database you cannot export, you do not have it yet.

Structured change control is the boring version of the answer, which is why firewall change management keeps outliving the compliance frameworks that come and go on top of it. A change that was requested, reviewed, approved, implemented, and recorded is defensible in front of any regulator, in any year, under any acronym.

If you want an outside read on whether your current evidence would hold up, the free NIS2 Readiness Check walks through what an assessor actually asks for.

About FwChange

FwChange is a Firewall change management methodology

Full Bio →FwChange Methodology
FW

FwChange

Firewall change management

Methodology and software for firewall change management, drawn from a large dataset of enterprise firewall migrations.