Pilot Launch Plan
Eleven national pilots across four waves
Eleven national pilots (§24) plus a first-year trajectory (§29, screen 16). Timelines are a working hypothesis, refined together with the institutions.
A principle shared by every pilot: none requires custom software development. Each uses off-the-shelf agentic environments (OpenAI Codex, Anthropic Claude Code, open-source equivalents); "deployment" = publishing configurations (instructions, roles, skills) + connecting data + training people. The technical launch of any pilot is possible on the day the agreement is signed.
1. Pilot passports
Passport format: target user · problem · minimum data · configuration · artifacts · key risks · success criteria · time to practical effect.
P-1. Workstation for a local self-government lawyer
- User: a lawyer in a hromada (community) executive committee — often 1–2 lawyers for the whole hromada.
- Problem: the broadest possible range of questions (leases, land, HR, procurement, citizen appeals) with no narrow specialisation and no mentors.
- Minimum data: legislation (open API), the hromada's model contracts, local acts.
- Configuration: baseline + skills "lease agreement", "reply to a citizen appeal", "review of a draft decision".
- Artifacts: baseline level; extended level for council/executive-committee decisions.
- Risks: weak on-site technical support; personal data in citizen appeals.
- Success criteria: time to prepare a standard reply ↓ ≥40%; 100% of replies with a source register; lawyer satisfaction.
- Time to effect: 4–6 weeks.
P-2. Assistant for analysing draft normative acts
- User: lawyers in the secretariats of Verkhovna Rada committees, ministries, and central executive bodies.
- Problem: a large flow of drafts; manual search for conflicts with current legislation and EU law.
- Minimum data: the Rada's API (draft laws, laws in force), texts of the relevant EU directives.
- Configuration: role "norm analyst"; skills "conflict map", "acquis compliance table".
- Artifacts: extended level: rules_matrix by norm, critique_pack listing the conflicts.
- Risks: political sensitivity of wording; responsibility for a missed conflict stays with the human.
- Success criteria: coverage of the acts reviewed ↑; identified conflicts confirmed by experts (precision ≥80% on reference cases).
- Time to effect: 6–8 weeks.
P-3. Judge's assistant for structuring case materials
- User: a judge and a judicial assistant (pilot — 1–2 courts, civil/commercial jurisdiction).
- Problem: caseload many times above the norm; hours spent building timelines of the case files and an evidence map.
- Minimum data: the materials of specific cases (isolated perimeter!), the USRCD for case law.
- Configuration: role "case structurer": timeline, fact_matrix, evidence map; no draft rulings at the start (to reduce the pilot's sensitivity).
- Artifacts: audit-grade level (high cost of error), local processing.
- Risks: maximum sensitivity; the perception that "AI is judging"; UJITS requirements.
- Success criteria: time to prepare a case for hearing ↓ ≥30%; zero confidentiality incidents; subjective usefulness per a survey of judges.
- Time to effect: 8–12 weeks. Partner: State Judicial Administration / National School of Judges + Pravo-Justice.
P-4. Assistant for a civil servant preparing legal replies
- User: the legal departments of central executive bodies and oblast state administrations.
- Problem: a flow of appeals and requests with tight deadlines; template work at a high cost of inaccuracy.
- Minimum data: legislation, the agency's base of previous replies (de-identified).
- Configuration: skills "reply to an appeal", "reply to a public-information request", with citation checking.
- Artifacts: baseline level + mandatory citation-check.
- Success criteria: reply turnaround ↓; share of replies sent back by a supervisor for rework ↓ ≥50%.
- Time to effect: 4–6 weeks.
P-5. A university workstation for students and teachers
- User: a law faculty (2–3 partner universities).
- Problem: graduates cannot work with AI professionally; universities have no methodologies.
- Minimum data: open data (USRCD, legislation) — no confidential layer needed.
- Configuration: educational: teaching cases, a "show your reasoning" mode, grading of artifacts.
- Artifacts: extended level as a teaching format (the student submits a package, not an essay).
- Success criteria: the course is added to the curriculum; ≥100 students; a bank of ≥20 teaching cases in an open repository.
- Time to effect: one semester.
P-6. Assistant for harmonising legislation with EU law
- User: the government's European-integration office, line ministries.
- Problem: screening is complete; ahead lies the systematic alignment of thousands of acts with the acquis, cluster by cluster.
- Minimum data: EUR-Lex (ELI, machine-readable), Ukrainian legislation, screening tables.
- Configuration: role "harmoniser": mapping directive↔law, compliance tables, tracking of acquis changes.
- Artifacts: extended; compliance tables are the key reusable asset.
- Success criteria: time to prepare a compliance table ↓ ≥50%; tables reused across agencies.
- Time to effect: 8–10 weeks. Funding: logically via the Ukraine Facility mechanisms.
P-7. A workstation for analysing case law
- User: analysts at the Supreme Court of Ukraine, the National School of Judges, law firms, research institutes.
- Problem: 120M+ decisions; manual analysis of samples is unrepresentative.
- Minimum data: USRCD datasets (already open, updated daily).
- Configuration: skills "case-law sampling", "map of positions", "detection of divergent case law".
- Artifacts: extended; source coverage is critical (each thesis → specific decisions).
- Success criteria: reproducibility of reviews (another analyst repeats the sample from queries.jsonl); identified divergences in case law are confirmed.
- Time to effect: 6–8 weeks.
P-8. Assistant for state legal services
- User: lawyers representing the state in court (Ministry of Justice, agencies).
- Problem: mass, uniform disputes; losses due to missed arguments and deadlines.
- Minimum data: USRCD (case law by category), the agency's de-identified case materials.
- Configuration: roles "dispute analyst" and "reviewer"; skill "map of the parties' arguments".
- Artifacts: extended + team exchange between regional lawyers (linked to P-11).
- Success criteria: consistency of positions within a category of disputes; time to prepare a response ↓ ≥40%.
- Time to effect: 8–12 weeks.
P-9. A tool for reviewing contracts and regulations
- User: lawyers at state-sector enterprises, municipal utilities, procurement officers.
- Problem: typical risks in contracts are discovered after signing.
- Minimum data: legislation, the institution's library of model contracts and checklists.
- Configuration: skill "contract review": a risk checklist + citation-check + red flags.
- Artifacts: baseline; the review report is a standardised artifact.
- Success criteria: checklist coverage 100%; risks found on a reference sample ≥85% against expert labelling.
- Time to effect: 3–4 weeks (the fastest pilot, a good first public case).
P-10. A system for training and certifying Legal AI mentors
- User: future mentors: advanced lawyers, legal-tech specialists, teachers.
- Problem: without a mentor layer, the other pilots do not scale.
- Minimum data: a curriculum, reference cases, an open repository.
- Configuration: a training track: setting up the workstation → formalising the process → the artifact standard → security.
- Artifacts: a capstone project — a real configuration for a specific institution.
- Success criteria: a first cohort of 20–30 certified mentors; each accompanies ≥1 deployment.
- Time to effect: 10–12 weeks per cohort.
P-11. A working group on full exchange of structured artifacts (the key pilot)
- User: one real team of 5–8 lawyers (an agency legal department or a partner law firm).
- Problem: to prove the team effect — the concept's central hypothesis (T14).
- Minimum data: the team's workflow; the artifact standard v0.1; a Git perimeter.
- Configuration: the full cycle: author/reviewer/supervisor roles, handoff gates, validators.
- Artifacts: all 8 types; three levels by task type.
- Risks: resistance to changing habits; over-complication of the standard.
- Success criteria (standard metrics): time to review ↓ ≥50%; unsupported claim rate → 0; handoff completeness ≥80%; source coverage ≥90%; the team voluntarily continues after the pilot.
- Time to effect: 12 weeks. The pilot's result is the principal evidential material for the whole concept.
2. Priority matrix
| Wave | Pilots | Rationale |
|---|---|---|
| Wave 1 (quick wins) | P-9, P-1, P-4 | Low sensitivity, fast measurable effect, public cases |
| Wave 2 (core of the evidence) | P-11, P-10, P-5 | The team hypothesis + the mentor layer + education |
| Wave 3 (institutional) | P-2, P-6, P-7, P-8 | Require agency partnerships |
| Wave 4 (maximum sensitivity) | P-3 | After security and trust have been worked out |
3. First-year trajectory
0–3 months. Open repository · Wave 1 configurations (5–10) · artifact standard v0.1 + validators · training materials · agreements with 3–5 pilot institutions · baseline measurements (how they work today).
3–6 months. Launch of Waves 1–2 · pilot P-11 · the first mentor cohort (P-10) · monthly metric snapshots · publication of interim cases on the landing site.
6–12 months. Waves 3–4 · artifact standard v0.2 based on feedback · a practice network (regular meetings of institutions) · a public metrics report · proposals for state-level institutionalisation (the role of a coordinating centre, certification).
Continuously: training, collecting feedback, updating the evidence base.
4. Resources and partnerships
- Core: 2–4 people (concept architect, configuration engineer, methodologist, pilot coordinator).
- Institutional partners: 1 university (P-5), 1 hromada (P-1), 1 agency (P-4/P-8), 1 team for P-11; for P-3 — the State Judicial Administration / National School of Judges in Wave 4.
- Support programmes: EU Project Pravo-Justice (justice), EU4DigitalUA / digital programmes (data), the Ukraine Facility mechanisms (harmonisation — P-6), university grants.
- Infrastructure: minimal — users' off-the-shelf agentic environments (commercial licences or open-source, with no development whatsoever) + Git + an open repository; an isolated perimeter only for P-3/P-8.
5. Stop criteria (an honest pilot)
A pilot is wound down or revised if: quality metrics fail to reach their thresholds in two consecutive snapshots · a confidentiality incident occurs · the standard's overhead eats the time saved (time to prepare ↑ over 25% with no compensating gain in time to review) · after training, the team does not use the tool voluntarily. A negative result is published exactly like a positive one — this is part of the project's scientific integrity.