In partnership with

Regular readers will know this site's short series stress-testing the legal AI skills published on Lawve by Larissa Meredith-Flister: her Opposing Counsel Review and Judicial First Impression skills were each run against a deliberately flawed skeleton argument, and the full outputs published with commentary. The premise of those articles was that the profession benefits when skills are tested in public rather than taken on trust.

This article applies the same premise, but this time against my own skill. I have published a skill of my own to Lawve, a client instruction schedule builder for civil litigation, and before doing so I put it through the marketplace's own audit tooling. This piece covers what the skill does, why it exists, what the audits found (including the things they caught that I had missed), and what changed as a result. As with the earlier articles, the embarrassing parts are left in: that is rather the point.

The problem: a client drowning in questions

The skill began with a real situation, described here in general terms. A client in a document-heavy dispute, entirely engaged and entirely willing, was simply overwhelmed. Not because the case did not matter to them, but because what we were asking of them had become impossible: two long witness statements from the other side, an advice from counsel, and a series of letters each asking for "comments" on the growing pile. Every request was reasonable on its own. Cumulatively they amounted to "please re-read everything and tell us everything", which is a task no anxious lay client can start, let alone finish.

Of course, as any litigation solicitor knows, this is a common problem and whilst we do all we can to support clients, we cannot actually answer their evidence for them.

Litigators do already have a tool for imposing discipline on sprawling factual disputes: the Scott Schedule, a table which forces each party to state its position on each item, one row at a time. The insight I had (if it even deserves the word given its simplicity) was that the same discipline works for taking instructions, provided the table is written to the client in plain English rather than to the court in pleading-speak. One row per issue. What we say, what they say (with a precise reference to where they say it), what evidence we already hold, and then specific, mostly closed questions: not "please comment on paragraph 14" but "did your husband speak to Mr Chudleigh in mid-February 2025, and what was said?". A priority column tells the client which four rows actually matter, so they can triage instead of freeze.

This is accompanied by a one-page covering note which explains that "cannot recall" is a proper answer and that returning the schedule in stages is fine.

The format worked. Over this weekend I turned it into a skill: a set of written instructions which any capable AI model can follow to read a case folder, extract the issues on which client input is genuinely needed, and produce the schedule and covering note as Word drafts for solicitor review. The skill file encodes the hard-won details: identify documents by their content because filenames lie, never ask the client for a document the firm already holds, flag every inconsistency in the papers to the solicitor instead of quietly smoothing it over, and stamp every page of every draft with a banner saying it is an AI-assisted draft which has not been approved or sent.

What it produces

The demonstration matter used for this article, Frayne v Kestrel Groundworks & Building Ltd, is entirely fictional: an invented barn-conversion dispute with duelling quotations, contested variations, a departed builder and a roof which may or may not have been undone by a storm. Run against the fictional papers, the skill produced a four-page landscape schedule of eight numbered rows and a one-page covering note. A representative row, verbatim:

Issue 3: How the builders came to leave the site.

Our position: Kestrel's workers did not arrive on 17 February 2025 and never returned, leaving the works incomplete (your 1st statement, para 6.1).

The other side's position: On 19 February 2025 your husband told Mr Chudleigh to "get your vans off my drive", which Kestrel treated as an instruction to leave, writing the same day to reserve its position (Mr Askey's 2nd statement, para 14; exhibit TA2, page 22).

Evidence we already hold: Both statements; Kestrel's letter of 19 February 2025 (exhibit TA2, page 22).

What we need from you: Did your husband speak to Mr Chudleigh in mid-February 2025? If so, on what date, and what was said, as best you can recall? Did you ask or authorise anyone to tell Kestrel to leave? "Cannot recall" is a perfectly proper answer.

Priority: High.

That row does precisely what is required. The two statements disagree about the date the builders left. The schedule does not resolve the disagreement or hide it: rather, the client is asked for her recollection, and a highlighted note above the table tells the reviewing solicitor to date-check the letter in the exhibit before anything is put in a court document.

The fictional papers were seeded with three such traps (a mislabelled exhibit bundle, a retention figure which differs between the statement and the pleading, and the conflicting departure dates), and the skill's job is to surface all three as solicitor-facing notes rather than to guess its way past them.

The "evidence we already hold" column deserves a word too, because it is the column that keeps the tool honest and easily verifiable. Everything relevant already on file must be listed there, which makes it structurally difficult to ask the client for something the firm already has. In the demonstration matter, the client is not asked to re-send her bank statements; she is asked to obtain statements for one closed account the firm does not hold, and the covering note offers a template letter to make that easy.

An example of the output, from the test scenario

The stress-test: running the marketplace's auditors against my own work

The earlier articles in this series tested other people's skills. Publishing my own imposes a different obligation, and Lawve's catalogue conveniently contains the instruments: skill-injection-defense, which audits skill packages for prompt injection, hidden instructions, unsafe scripts, credential exposure and exfiltration paths before they are trusted or published, and epistemic-fault-line-audit, which audits legal AI instructions for fluent but unsupported reasoning, overconfidence, hidden assumptions and missing human-review gates.

Both are by Ignacio Adrián Lerer, whose auditing skills are, in this writer's view, among the most quietly valuable things on the marketplace.

The security audit came back clean: a verdict of "approve", no injection, no hidden Unicode, no network calls, no credentials, with three hygiene recommendations. All three were adopted.

The epistemic audit was the interesting one, and it did not come back clean. Verdict: "material fault lines".

Two findings were interesting, because both were correct and neither had occurred to me.

The first: the skill's verification steps checked only what was in the draft, never what had been left out. The workflow validated the file, rendered every page, confirmed the numbering and re-checked every reference, and at no point asked whether an issue in the papers had silently failed to become a row. As the auditor put it: "its verification checks only what the draft contains, never what it omits, so the reviewing solicitor is shown a polished table with no way to see the issues the model silently dropped". For a tool whose whole purpose is to be exhaustive on behalf of an overwhelmed client, that is a critical structural blind spot, and the fix is now in the published skill. Before drafting, the model must build a source-to-row map (every disputed statement paragraph, every "we need instructions on" passage in counsel's advice, every unanswered request in correspondence), and every mapped item must end up either as a row or in an "issues considered and excluded" appendix, with reasons, inside the document itself. The reviewing solicitor now sees the exclusions on the same pages as the inclusions.

The second finding was more lawyerly. The skill mandates mostly closed questions, and its table shows the client the other side's position before asking for her recollection. If the answers later feed a trial witness statement, that is uncomfortably close to the practice which Practice Direction 57AC exists to restrain: putting an account to a witness and inviting agreement or disagreement on important disputed matters. The original skill file never mentioned it. The published version now carries an express PD 57AC caveat: where answers may feed witness evidence, the solicitor should consider generalising or removing the "other side's position" column for contested recollection issues, and should keep a record of what documents the client was shown. The tension between "specific closed questions an overwhelmed client can actually answer" and "open recollection evidence the court will accept" is real, and the honest position is that the schedule is an instructions-gathering tool first, with the witness-evidence implications flagged to the human whose job it is to manage them.

The audit found nine further fault lines of varying weight, most of which produced changes. The verification step now requires re-reading each cited paragraph to confirm the summary is a fair paraphrase (existence of a document is not faithfulness to it), priorities are now checked against limitation-sensitive and decision-blocking issues, and the skill now instructs the model to flag documents referenced in the papers but absent from the folder.

Readers of the earlier articles will recognise the shape of this experience. Running a well-designed audit skill against a piece of work is cheap, fast, and mildly bruising, and the work is better for it. That was true when the work was a skeleton argument. It turns out to be equally true when the work is a skill.

Talk to your AI tools the way you'd talk to a colleague.

You don't send a colleague a three-word brief. You explain the context, the constraints, what you've already tried. But typing all that into ChatGPT takes forever — so you don't.

Wispr Flow lets you speak your prompts instead. Talk through your thinking naturally and get clean, paste-ready text. No filler words. No cleanup. Just detailed prompts that actually get you useful answers on the first try.

Millions of users worldwide. Works system-wide on Mac, Windows, and iPhone.

The licence, in plain terms

The skill is published under the Apache 2.0 licence, the same licence this site praised when examining Meredith-Flister's skills, and the choice is deliberate for the same reasons. UK firms can download the skill, use it on paying client matters, and modify it to fit their own precedents and house style, all without permission or payment. Attribution must be preserved, modified files must be marked, and the licence's express patent grant removes a risk which makes some firms' open-source policies twitch about barer licences such as MIT. Nothing a firm builds on top of it (bespoke columns, integration with a case management system, a version tuned to construction disputes) has to be published back to the world, though this author would be delighted if improvements were.

One caution belongs here rather than in a footnote. The skill instructs an AI model to read an entire litigation folder, including counsel's advice. Whether that is permissible in your practice depends on your firm's AI and confidentiality policy and on the tool you run it in, and the published skill now says so on its face. It also says, in three separate places, what the banner on every generated page says: the output is a draft for solicitor review, and no part of it is legal advice.

How to get it, and an invitation

The skill is available on Lawve here. The package is the instruction file, a complete working build script, and example data from the fictional Frayne matter; it runs in any AI environment which can follow a skill file and generate Word documents, without vendor lock-in.

The invitation is the same one this series has extended to other authors' work: download it, run it against a synthetic or anonymised matter, and try to break it. The audits above caught real problems, and I have no illusions that they caught all of them. If a row misstates a position, if a question leads where it should ask, if the coverage map misses a category of source material you rely on, tell me, and the next version will be better for it.

How did we do?

If you try the skill, hit reply and tell me what happened, including and especially if it embarrassed itself. I read every email.

Thanks for reading,

Serhan, UK Legal AI Brief

Disclaimer

Guidance and news only. Not legal advice. Always use AI tools safely and in line with best practice.

Keep Reading