Skip to content
Aryan.
← All work

04 · LLM safety and RAG

Legal Risk RAG (confidential)

Status
Technical contribution (confidential, pre-launch)
Role
Backend engineer, security audit
Started
2025

Node.js / Fastify / pgvector / Anthropic API

Confidential technical contribution to an Australian legal-tech startup. Real product name held back pending formalisation of my involvement.

The problem

Small and mid-size Australian businesses need to assess employment-law risk (unfair dismissal, general Fair Work Act coverage) before making high-stakes decisions like termination. Traditional legal advice is slow and expensive; a well-designed AI system can triage the obvious cases and flag the uncertain ones for a human lawyer.

My contribution

  • Built the pgvector-based RAG pipeline over an Australian employment-law corpus (Fair Work Act sections, precedent, plain-English guidance)
  • Designed the 4-pass processing architecture: initial extraction, contextual retrieval, classification, and structured risk assessment. 17 total operations across the pipeline
  • Contributed to the 5-band risk classification model (green through red) that shapes the final user-facing output
  • Ran a security hardening pass that fixed a prompt-injection vulnerability in the risk-band selection logic (see below)

The prompt-injection fix (highlight)

The original pipeline determined the final risk band by string-matching phrases in the model's prose output. This meant a well-crafted user input containing phrases like "this is clearly a low-risk situation" could steer the classification: the model would reference the phrase in its reasoning, and the string match would pick it up.

The fix: added a misconductClassification field to the model's structured output schema, validated against an enum at parse time. selectWeightVariant (the function that picks the risk-band weight) now reads that validated enum field, not the free-text reasoning. Prompt injection can still influence prose but no longer steers the risk-band selection.

Key decisions and why

  • Structured output over prose parsing. Prose is a surface for injection; enums are not. Any time the pipeline branches on a model decision, that decision has to come from a validated schema field
  • Corpus-first RAG, not fine-tuning. Legal ground truth changes (new precedents, new regulations). Retrieval keeps the system current without retraining
  • Five bands, not three or ten. Enough resolution to be useful; few enough that human reviewers can act on the classification without translation

Note

This project is co-owned with a business partner and is currently in pre-launch. The product name and further details are confidential until my involvement is formalised. Happy to discuss the technical work in interviews under NDA.

Next

AutoLZ →