Harvey put out three procurement-facing announcements in four days: an AIUC-1 safety certification, a Goldman Sachs and J.P. Morgan investment, a new chief revenue officer. What they establish for a legal team is a narrower question than the framing suggests. OpenAI's next model quietly solved ten open problems in mathematics, with proofs a machine can check line by line. On both sides of the Atlantic, regulators let the transparency paperwork land on schedule and pushed the substantive high-risk rules into next year. And Crosby, the AI-native law firm, decided the way to let its agents work without a lawyer signing off is to insure them like one.
Harvey made three announcements in four days: an AIUC-1 certification, growth capital from two bank investment arms, and a new chief revenue officer. All three are aimed at the same audience, the enterprise buyers who decide whether software gets trusted with real work, and each is worth reading for what it establishes rather than for how it was framed.
On August 4, Harvey became the first legal AI company certified against AIUC-1, a standard for AI agent security, safety and reliability that has existed for barely a year. The auditor, Schellman, ran more than 3,000 tests across six domains: data and privacy, security, safety, reliability, accountability and societal impact. Harvey reports zero critical failures. The framework maps onto standards enterprise security teams already use to vet vendors, including ISO 42001, the EU AI Act and NIST's AI risk management framework, which is precisely what makes it useful as a sales document: it pre-answers procurement's security questionnaire. What it can't yet claim is a track record. No one knows whether an AIUC-1 pass predicts how an agent behaves in production, because the standard is too new for anyone to have checked.
The certification landed a week after Growth Equity at Goldman Sachs Alternatives and J.P. Morgan's Growth Equity Partners took a stake of undisclosed size, on the back of a quarter in which Harvey says it added more than $100M in ARR. On August 5, Harvey named Steve Zad, previously of Rubrik, as its first chief revenue officer. These are the ordinary mechanics of a growth-stage software company preparing to sell more software. They say a lot about Harvey's go-to-market ambitions and nothing about whether the work a legal team hands over comes back done.
A safety certification tells a buyer the tool won't misbehave. It doesn't tell them the work got done to their standard, or that it left the desk of someone too expensive to be doing it. Those are the questions a services relationship has to answer, and no compliance framework answers them by proxy.
Three labs shipped three different kinds of news this week, and the gap between them says more about where the industry actually is than any single release does.
On August 1, OpenAI published something unusual for a model that hasn't shipped: proof. An internal version of its next model, Astra, resolved or made substantial progress on ten open problems across eight fields of mathematics and theoretical computer science, including a 27-year-old question in group theory and two long-standing Erdős problems in combinatorics. What made the announcement notable wasn't the difficulty of the problems. It was the form of the answer. Every result came with a machine-checkable Lean 4 certificate, published openly on GitHub, so a mathematician doesn't have to take OpenAI's word that a proof holds. They can run the checker themselves. Fields medallist Timothy Gowers said he would recommend one of the results for a top journal without hesitation. The whole exercise cost roughly $2,000 in compute.
Meanwhile, Google's most anticipated model of the year still hasn't shipped. Gemini 3.5 Pro was announced in May, targeted for June, then July, then a specific mid-July date that also passed. As of this week it's rumoured for August 12, unconfirmed by Google, which would make it the fourth slipped date for a single release. xAI is running the opposite playbook. Grok 4.5 launched three weeks ago, and Elon Musk says Grok 4.6 is coming within days, a major-version cadence measured in weeks rather than the months every other lab treats as normal.
If your legal AI vendor is built on someone else's model, which of these three release patterns are you actually exposed to? A verified proof, a missed date, or a version bump you didn't ask for? Almost no legal AI buyer asks their vendor this question, and almost every vendor's honest answer changed at least once this year.
On August 2, the EU and California both let an AI transparency deadline land exactly on schedule, while quietly waving through the harder substantive rules that were supposed to arrive alongside it.
In the EU, Article 50 of the AI Act, the provision requiring disclosure when content is AI-generated and labelling for deepfakes and chatbots, became enforceable across the bloc precisely when the original text said it would. It wasn't touched by June's Digital Omnibus. What the Omnibus did move is Annex III, the high-risk system obligations covering things like biometrics, employment screening and credit scoring. Those now apply from December 2027 for standalone systems and August 2028 for high-risk systems embedded in products, more than a year later than the law originally promised. The Act's most consequential paperwork requirement showed up on time. Its most consequential compliance burden did not.
California ran an almost identical play, on the identical day, on purpose. SB 942, the state's AI Transparency Act, requires large generative AI providers to offer a free content-detection tool and embed provenance watermarks in AI-generated images, audio and video. Its operative date was expressly amended to land on August 2, the same day as the EU's transparency rules. Two jurisdictions that otherwise regulate AI on almost nothing in common chose the same date for the same category of obligation: tell people what's AI-generated. Neither one moved the harder question, what an AI system is allowed to decide on its own, at the same speed.
| Obligation | EU AI Act | California SB 942 |
|---|---|---|
| Took effect Aug 2, 2026 | Article 50 disclosure and labelling duties | Free AI-content detection tool; visible disclosure option |
| Pushed later | Annex III high-risk system obligations | Embedded provenance watermarking on images, audio, video |
| New date | Dec 2027 (standalone); Aug 2028 (embedded) | Phased through 2028 |
Disclosure regulation is cheap to comply with and politically easy to enforce on schedule. Substantive regulation, telling a system what it can't decide on its own, is neither. For a legal team building an AI governance programme, that gap is the actual risk map. The labelling rules are real now. The rules that would have forced a hard look at autonomous decision-making just bought everyone another year and a half.
Legal AI's hardest unresolved question isn't whether agents can do the work. It's who answers for it when they don't. This week produced two very different answers, from two very different layers of the market.
Crosby, the venture-backed AI law firm that reviews contracts in under an hour, announced it will buy professional liability insurance for its own agents, so they can complete legal work without a human lawyer signing off first. Founder Ryan Daniels was direct about where this is heading. Lawyer review of AI output, he said, won't be necessary in the future. It remains, for now, the standard the rest of the industry still runs on. The logic is coherent on its own terms. If the firm is financially on the hook when an agent gets it wrong, in theory it has every incentive to make sure it doesn't. But it moves the mechanism of trust from verification to indemnity. Nobody checks the work before it goes out. Somebody just pays for it afterwards.
Further down the stack, the software layer kept moving at its usual pace, unbothered by the accountability question. Wordsmith closed a $14M extension to its Series B, taking its total raised to roughly $114M. Ironclad shipped new agents that scan supplier contracts for buried rebates and compliance obligations and route them straight into SAP. Both are incremental tool improvements. Neither one changes who is answerable when the output is wrong, because in both cases a person is still expected to check it before it matters.
The second model is slower to describe and harder to fit in a press release. It's also the only one of the two where the client, not an underwriter, still owns the standard the work is held to.
Every story this week is, underneath, the same story. The industry is racing to prove AI can be trusted with real stakes faster than anyone has agreed on what that trust should actually rest on.
Every certification, insurance policy and disclosure rule this week was, in its own way, an attempt to answer the question the market still hasn't settled: who is responsible when AI does real legal work. None of them touch the problem underneath it, which is that inexpensive, repetitive work is still being done by the most expensive people in the building, and a badge on the tool doing it faster doesn't change that. Harvey's AIUC-1 certification says the agent is safe to run. It says nothing about whether the NDA left a general counsel's desk.
Crosby's answer, insuring the agent so it can work unsupervised, is at least honest about the stakes. But it transfers risk rather than removing it, and it keeps the work inside a vendor's four walls rather than the client's own. Insourcing enterprise legal work to supervised agents is a different bet: the client's own playbooks, the client's own escalation rules, and a supervision layer, run by the client's lawyers, a partner firm, or Flank's own team, that catches the exception before it ever reaches the business. Tools are a commodity. A certified tool is still a commodity. Outcomes, work that leaves the desk and doesn't come back, are not, and nothing published this week claimed to deliver one.