The White House accused Moonshot AI of covertly distilling Anthropic's Fable to build Kimi K3 — a launch so popular Moonshot had to pause new subscriptions within 48 hours — while Google quietly shipped three Gemini models and left its long-teased flagship unshipped. Crowell & Moring published six months of firmwide Legora usage data, and two federal judges took opposite approaches to AI-hallucinated filings in the same week. And with nine days left before the EU AI Act's Article 50 transparency rules bind regardless of the high-risk deferral, most legal teams are bracing for the wrong deadline.
Model distillation, feeding a stronger model's outputs to a weaker one so it learns to imitate the stronger one, is a normal, often licensed, technique in AI development. On July 23, White House Office of Science and Technology Policy director Michael Kratsios accused Moonshot AI, the Beijing-based lab behind Kimi K3, of doing it covertly: running the copying through a purpose-built internal system and rotating access routes to stay hidden while pulling outputs from Anthropic's Fable model. Kratsios also said Moonshot obtained access to export-controlled, high-end Nvidia AI servers to train the model, hardware that is barred for export to Chinese entities without a license. Anthropic itself had flagged unusual activity back in February, tracing 3.4 million Claude exchanges to the Chinese startup. Moonshot has not responded to the allegations. No penalties have been announced. And the punchline is almost absurd: Kimi K3's own public launch a few days earlier was so popular that Moonshot had to pause new subscriptions within 48 hours, "the love," as the company put it on social media, outstripping its GPU capacity.
Every enterprise legal AI contract has a clause, somewhere, about what the vendor can and can't do with your firm's inputs and outputs. This dispute is the first widely reported case of a frontier lab allegedly doing exactly that to a rival, at industrial scale, in violation of usage terms. If a well-resourced state-linked lab can allegedly do this to Anthropic and still ship a product before anyone can prove or stop it, what confidence should an enterprise buyer have that its own outputs, contract clauses, negotiating positions, internal analysis, aren't being harvested by whatever model sits underneath a vendor's product today? Provenance of training data is becoming a genuine legal due-diligence question, not a theoretical one.
This story is a reminder of something this briefing keeps returning to from different angles: the intelligence layer underneath every legal AI product is rented, contested, and increasingly adversarial. Whether or not the specific allegations against Moonshot hold up, the incentive they describe, that outputs are valuable enough to steal at industrial scale, applies just as much to the outputs your legal team generates inside whatever AI tool it uses today. Flank's model doesn't ask a client to trust a single model vendor's training practices as a matter of faith. Supervision sits above the model layer: a human reviews every agent output before it reaches the business, and the workflow, not any one underlying model, is what the client is actually buying. When the model layer is this volatile, the thing worth building trust around is the review step, not the vendor's word about what its model was trained on.
Google shipped three new Gemini models on July 21: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a narrow, security-focused variant called Gemini 3.5 Flash Cyber. 3.6 Flash cuts output token pricing by roughly 17% against its predecessor and posts double-digit gains on coding and reasoning benchmarks, but Gemini 3.5 Pro, the flagship this briefing has now heard teased across two editions, following a full architectural rebuild after the original base model was scrapped, still has not shipped. Meanwhile Moonshot's Kimi K3, a 2.8-trillion-parameter open-weight model, launched the same day and proved so popular that Moonshot paused new subscriptions within 48 hours to protect existing users' access, before the distillation allegations above had even surfaced.
Three models, three different states, shipped, promised, and overwhelmed, inside a single week, and almost none of it is visible to the person typing into a legal AI product's chat window. Does your legal AI vendor tell you when the model underneath your workflow changes, and does your audit trail record which model produced a given output? A flat interface hides model churn by design. A governed workflow has to surface it, because the model that drafted last month's NDA redline may not be the model that drafts next month's.
The political fight this briefing covered across April and May, over whether the EU AI Act's high-risk compliance deadline would hold at August 2026 or slip, is over: the Digital Omnibus on AI was signed into law on July 8, formally deferring high-risk obligations to December 2026 and December 2027 in stages. But Article 50, the AI Act's transparency regime, was not touched by that deferral and becomes enforceable on schedule on August 2, 2026: mandatory disclosure when a user is interacting with an AI system, labelling of AI-generated synthetic content, and disclosure of deepfakes. Unlike the high-risk provisions, Article 50 isn't limited to systems classified as high-risk. It applies to any business using generative AI to produce content that reaches an end user, a considerably wider net than the compliance conversation of the last four months has focused on.
The practical risk is that legal and compliance teams who spent the spring tracking the high-risk deferral fight now read the Omnibus headline, "deadlines pushed back," and stand down, when the obligation that actually lands in nine days was never part of that fight. Chatbot disclosure, AI-content labelling, and deepfake marking apply whether or not a given tool is high-risk, which in practice covers most of what a legal team's own AI vendors are producing for it today: drafted correspondence, summarised research, generated first-pass documents.
This is a small, specific illustration of a pattern this briefing keeps documenting: regulatory complexity doesn't resolve into a single deadline a legal team can put on a calendar and forget. It resolves into a set of overlapping obligations that a governance layer has to track continuously, not a compliance program that gets built once. Flank's supervision model exists for exactly this reason: every agent workflow runs through human review before output leaves the system, which means disclosure, labelling, and audit-trail obligations are met as a structural property of how the work gets done, not as a separate compliance exercise bolted on after the fact once the deadline is finally clear.
Crowell & Moring piloted Legora in late 2025 and launched it firmwide in January 2026. On July 22, the firm published six-month results: more than 2 million platform interactions, 82% of the entire firm, including 91% of its more than 700 attorneys, now counted as users, and nearly 70% of those attorneys using the platform at least weekly. That last figure is the one that matters most. Adoption numbers are easy to publish and easy to inflate with one-time logins; weekly active usage at that depth, across a firm of that size, is a different claim entirely, evidence of the tool being embedded in how the work actually gets done rather than sitting unused after a launch announcement.
An associate in the firm's Patents Group described the platform as having "accelerated my learning curve and, therefore, value to clients in ways I didn't anticipate." The firm also cited a recent trial where the team used the platform to prepare more thoroughly and respond more nimbly than an opposing team with more people on it. Both are anecdotes rather than measured outcomes, and Crowell & Moring didn't publish a cost or time-saved figure to sit alongside the usage data, which is itself notable given how squarely last week's ROI-measurement gap sits over this exact kind of announcement.
Weekly-active-usage data like Crowell & Moring's is genuinely useful, and rare. But usage depth answers a different question than the one Axiom's survey found 83% of in-house teams can't answer: not "are people using the tool," but "does the work cost less or get done better because of it." If your own legal AI vendor can't produce Crowell & Moring's kind of usage number, and your own team can't produce a cost or outcome number to go with it, which of the two gaps is actually the harder one to close?
This briefing has spent months tracking an escalation: sanctions climbing from four-figure fines toward $15,000-per-attorney penalties, bar suspensions, and disqualifications. This week ran in the other direction in two separate cases. On July 17, Michigan federal judge Hala Jarbou caught a DOJ brief citing Taylor v. Hott, a Sixth Circuit case that does not exist, and warned the government that hallucinated law in federal filings is unacceptable, but declined to sanction. Around the same time, Kentucky federal judge Thomas Cullen declined to sanction attorney Thomas Guyer over a brief filled with AI-generated misquotes and incorrect citations, finding Guyer had "owned the mistake," had no history of misconduct, and was, in his own lawyer's words, "incredibly remorseful." Cullen was explicit that a warning could serve as "sufficient deterrent" on a clean record, while still insisting that generative AI's growing role in practice, its becoming the "new normal," doesn't relax a lawyer's duty to take reasonable measures to verify what gets filed.
Neither case suggests courts are going soft on the underlying problem. Both judges were explicit that verification remains the lawyer's job regardless of the tool. What they introduced is a distinction between the mistake and the response to being caught: remorse, a clean record, and taking ownership now appear to buy real leniency, in a way that a pattern of hallucinations, or a lawyer who fights the finding, does not.
The distinction these two rulings are drawing, first-offense contrition versus a pattern of disregard, is itself a supervision question, not a sanctions question. A firm that can show a court it has a real verification step built into how AI-assisted work gets produced and checked has a materially different story to tell than a firm that can only say a lawyer was sorry after the fact. That is the entire premise behind building supervision as core product rather than as a policy memo: every Flank agent workflow runs through human review before anything leaves the system, so the question a court is now implicitly asking, "was there a real check here, or just good intentions", has an answer that doesn't depend on how convincingly remorseful anyone sounds afterward.
A frontier-model dispute over stolen intelligence, three models moving in three different directions in a single week, a compliance deadline hiding behind a headline about deferral, a law firm's usage numbers with no outcome numbers attached, and two judges quietly rewriting how much mercy a hallucination earns: none of these stories were coordinated, and all of them describe the same market working through the same unresolved question. Intelligence keeps getting more contested and more volatile at the same time it gets cheaper and more available. Oversight keeps arriving in pieces, a transparency article here, a judicial distinction there, rather than as one governed system. And the gap this briefing keeps finding, between using AI and being able to prove what it did, hasn't moved.
Every story in this week's briefing is a variation on the same structural gap: inexpensive work is being done by expensive resources, and neither cheaper intelligence, nor a firm's usage numbers, nor a court's evolving sense of how much mercy a mistake deserves, closes it on its own. A frontier lab allegedly stealing another's outputs to build a faster model doesn't change who reviews the work before it reaches a client. A compliance deadline hiding behind a deferral headline doesn't change whether a legal team can actually show a regulator what its AI systems did. And a law firm publishing real, granular adoption numbers, genuinely rare, still isn't the same claim as being able to show the work costs less or gets done to a verified standard.
Flank's answer to all of it is structural, not a bet on any one model, vendor, or court's mood this quarter: tools are a commodity, outcomes are not. We compete for the budget line where routine legal work already sits, the outside counsel, ALSP, and internal headcount spend that dwarfs any software line, not just a new tool to add to it. Agents execute the work under your playbooks, a human reviews every output before it leaves the system, and you don't pay until the first workflow is live in production. Whichever model is running underneath any given vendor this week, whichever deadline just quietly bound, whichever judge just drew a new line on mercy, the one part of this week's news that doesn't change is that outsourcing the work to supervised agents is what turns "we use AI" into an answer a court, a regulator, or a CFO can actually verify.