Case building¶
The Case shell is the per-matter workspace: documents and their OCR text,
the Google Drive mirror that feeds them, highlights, the timeline of facts,
witnesses, labels, saved case law, and search inside the matter. It is one
Django app, apps/case, with one sub-package per tab, plus apps/drive
(the mirror) and apps/search (the sidebar search across matters). The
AI tab is documented on its own pages; this page only links to them.
Saved case law is a view of the AI tab in the interface but its code and
model live here.
Where the code is¶
| Path | What it holds |
|---|---|
apps/case/views.py, apps/case/urls.py |
The shell: tab switching, matter selection, every case: route |
apps/case/models.py |
Label, Document, Highlight, Fact, Witness, CaseLaw |
apps/case/documents/ |
Documents tab, upload, duplicate detection, mbox import, OCR task, viewer, the Drive Folder modal |
apps/drive/ |
The Drive mirror: changes feed, folder mappings, upsert, tombstones |
apps/case/highlights/ |
Highlights tab and the viewer's highlight endpoints |
apps/case/facts/, apps/case/witnesses/ |
Timeline and Witnesses tabs, the Timeline PDF |
apps/case/labels/ |
Labels tab and the apply-labels modal used by every tab |
apps/case/caselaws/ |
Saved cases (the AI tab's Case Law view), the case-law viewer, the summary task |
apps/case/courtlistener.py, courtlistener_throttle.py, jurisdictions.py |
The CourtListener client (citation lookup, cluster and opinion fetches, opinion search); the process-wide request throttle; the state court lists |
apps/case/search/, apps/case/search_config.py |
Search inside a matter; watson registration |
apps/search/ |
The sidebar search across matters, contacts, intakes and notes |
apps/accounts/middleware.py, apps/accounts/access.py |
Matter membership enforced from the URL |
templates/case/ |
One folder per tab; viewer.html and caselaw-viewer.html are standalone pages |
static/pdfjs/ |
pdf.js, installed by scripts/update_pdfjs.py; web/viewer.html is customised |
Data model¶
Every model carries AuditMixin (created_by, updated_by, timestamps)
and simple_history. Importance is one scale everywhere, 1 (Lowest) to 7
(Highest), 4 by default; apps/case/highlights/importance.py is the one
definition and parse_importance() the one parser.
Document:matter(CASCADE),category(Correspondence, Discovery, Evidence, Record),proceeding(SET_NULL),file,labels, the OCR fields (ocr_status,ocr_text,page_count,ocr_pages_done),search_vector,importance,ai_context(auto, always, never),summary, the fingerprints (content_hash,page_fingerprint) and the Drive provenance fields (drive_file_id, unique;drive_path;drive_modified;drive_mapping, SET_NULL).save()runsfull_clean(), enforces the proceeding rule and fingerprints a file that has no hash yet. History excludesocr_textandsearch_vector: a status-only save was snapshotting megabytes of text per write.Label:matternullable; null means a global label. Unique on (matter,name).Highlight: exactly one ofdocumentorcaselaw, enforced by the check constrainthighlight_has_one_source. Document highlights carrypage_numberandcoordinates(PDF rectangles); case-law highlights carrychar_offset. Six colours,labels,witnesses.Fact:matter,date,time,description(150 characters), optional rowcolor,documentsandhighlightsas sources,labels.Witness:matter, contact fields,alignment(friendly, neutral, hostile),knowledge.CaseLaw: a CourtListener cluster saved to a matter, unique on (matter,cluster_id).textandhtmlare legacy and no longer written; the opinion is fetched on demand.- Drive:
DriveFolderMapping(one top-level subfolder of a matter's Drive folder, itscategoryandproceeding; the check constraintdrive_folder_mapping_proceeding_rulemirrors the document rule),DriveRecordTombstone,DriveMatterState(the badge's cached state) and the single-rowDriveSyncState(the changes-feed cursor). The matter's folder itself isMatter.drive_folderanddrive_folder_id.
The proceeding rule lives in Document.save(): a document with a
proceeding is Record unless it is Discovery. Every path that writes
category and proceeding together respects it (the row menus refuse with a
toast, the form clears the proceeding, the Drive mapping writes the pair
through apply_mapping_fields()), but a bulk update() bypasses save()
and must keep the pair consistent itself.
How it works¶
The shell¶
templates/matters/matter-base.html holds a #mode-content block that the
mode pills switch between the matter's Detail view (matters:mode-content)
and the Case view (case:mode-content). Inside Case, case-nav.html is
the tab bar; each tab link GETs case:tab-content into .detail-body
and pushes the tab's own index URL. views.tab_content stores the tab in
the session (set_last_tab, keyed by matter), and _get_case_tab_data()
dispatches to the tab's get_*_data() function, which the tab's own
index and list views also call, so a full page load and an HTMX swap
build the same context. VALID_TABS in apps/case/views.py is the list a
stored tab is checked against, and tab_available(user, tab) drops the
ai tab when no AI provider is configured and caselaws without a
CourtListener token or perm_research, so a remembered tab that has
gone falls back to Documents. case-nav.html shows the AI tab only with
ai_enabled, and otherwise a Case Law tab when caselaw_available (see
AI is optional).
case-content.html and case-tab-content.html wrap every tab in a div
that re-fetches its list on a *Changed event from body
(documentsChanged, highlightsChanged, and so on); mutating views
answer 204 with that HX-Trigger rather than rendering. Filters, sort
keys, pagination and multi-select state all live in the session under
matter-scoped keys from get_session_key() (see
Session state).
Documents¶
documents_add accepts .pdf and .mbox only. For a PDF it fingerprints
the upload (fingerprint.py: SHA-256 of the bytes, and for PDFs a digest
of the page content streams and XObject data that ignores metadata), looks
for matches across every matter the user may see, and renders the form
back with the matches once; a second Submit with duplicate_ok uploads
anyway. The save is two-phase (row first, for the primary key the storage
path needs), then _file_under_chosen_matter() honours the form's Matter
field, _drop_other_matters_labels() strips another matter's labels, and
OCR is queued. An .mbox goes through mbox.py: parse_mbox_file()
reads the first message, generate_email_pdf() renders
templates/case/documents/email_pdf.html with WeasyPrint through a URL
fetcher that only allows the template's own font hosts, and the document's
date and description come from the email.
documents_edit refuses a new file on a Drive-synced document and a
matter change on one (DRIVE_MOVE_MESSAGE); a move of an uploaded
document copies the file to the new matter's path. Bulk actions read the
selection from the session and pass it through _selected_on_matter(),
which keeps only ids on the matter in the URL, because the membership
check covered that matter and not the selection. bulk_documents_matter
resolves the target through target_matter_for_user() because the target
arrives in the POST body.
download_document serves the file as an attachment under the document's
name (date-prefixed for Record documents); serve_document streams it
inline for the viewer, so the browser never needs the storage bucket's
origin. access.py holds open_matters_for_user(), the choice list for
the Matter field: the user's Open matters plus the record's own matter
whatever its status, because a select with no blank choice would otherwise
move the record on Submit.
OCR¶
tasks.process_document_ocr (Django-Q, group ocr_processing) reads the
file, extracts any text layer with pypdf, and if that is over
TEXT_THRESHOLD (500 characters) stores it with status extracted.
Otherwise it runs ocrmypdf with force_ocr=True, replaces the stored
file with the OCR'd PDF, and stores the result as completed. A
progress-bar plugin (ocr_progress.DatabaseProgressBar) writes
ocr_pages_done so the badge can show pages. Afterwards the search vector
is set from the first 900k characters (Postgres rejects a tsvector input
over 1MB) and, when ai_enabled(), generate_document_summary is queued. Failures set failed
and ocr_error; retry_ocr re-queues a failed document, accept_ocr
promotes extracted to completed, and force_ocr re-runs with
force=True for a document whose text layer was wrong. The badge
(templates/case/documents/ocr-badge.html) re-fetches itself with HTMX
while the status is pending or processing.
The viewer¶
document_viewer renders templates/case/viewer.html: a standalone page
with pdf.js in an iframe (static/pdfjs/web/viewer.html, loading
case:serve) and a highlights sidebar. The two halves talk by
postMessage (viewerReady, selectionChanged, scrollToHighlight,
setHighlights); a new highlight POSTs to add_highlight and the card is
built client-side. scripts/update_pdfjs.py installs a pdf.js release
from npm into static/pdfjs/ and preserves the customised
web/viewer.html across the update; web/viewer-custom.css is not
touched either. Highlights reach the page's script through
json_for_script(), which escapes what HTML would read, because a
highlight's text is whatever was selected in a PDF from outside the firm.
The Drive mirror¶
Setup is the Drive Folder modal on the Documents tab
(apps/case/documents/drive.py): pick the matter's folder under the Drive
root (DRIVE_NOTES_ROOT), then map its top-level subfolders to a category
and, for Record and Discovery, a proceeding. drive_folder_save validates
rows with mappings.normalize_rule(), writes DriveFolderMapping rows
keyed by Drive folder id (so renames do not break the link), runs
backfill_documents() over the documents a changed mapping already
brought in, and queues records.resync_mapping_by_id for each changed
row. Unmapping a row deletes the mapping and leaves its documents as they
are. suggest_mapping() pre-fills a category from folder names such as
"Corr" or "Record - Appeal"; Evidence is never suggested, on purpose.
Sync runs every minute (drive-sync, apps.drive.google.scheduled_sync)
over the Drive Changes API with a cursor in DriveSyncState, and nightly
(drive-sync-nightly-full) as a full crawl that also refreshes folder
names, flags missing folders and records each matter's unmapped
subfolders for the badge. _process_change() walks a file up to the root,
finds the nearest mapped ancestor with resolve_mapping(), and hands PDFs
to records.ingest_pdf(). That function is the upsert: a tombstoned id is
skipped; an unchanged modifiedTime only refreshes provenance (and strips
another matter's labels if the file moved between matters); a changed file
has its bytes replaced, fingerprints recomputed and OCR reset and
re-queued; a new file first looks for a hand-uploaded twin to adopt
(_adopt_candidate: same matter, name, date and size) before creating a
row. User-set metadata is never touched after the first ingest, and
category and proceeding follow the mapping only when the file arrives
through a different mapping than before.
The contract, stated in apps/drive/records.py: the mirror is
append-only (a deletion, trash or move out of Drive never removes a
document), PDFs only, and read-only on Drive (the application
never writes to Drive). A document deleted in the application writes a
DriveRecordTombstone through the pre_delete signal in
apps/drive/signals.py, so the file does not come back on the next pass.
Closing a matter clears its Drive link and mappings
(Matter.save()); the synced documents stay. restore_drive_documents
re-downloads documents whose file is missing from storage, a repair for a
window when uploads went to local disk (see
Prod storage).
Highlights¶
A highlight is made in one of the two viewers. The Highlights tab
(get_highlights_data) lists both kinds for the matter, filters by
document, witness, keyword, importance and source type, and sorts by a
key validated with with_valid_sort(). highlight_link is the View
Source menu item: it redirects to the document viewer at the page and
highlight, or to the case-law viewer. HighlightForm shows
paragraph_number for document highlights and page_number for case-law
ones; Highlight.citation builds a Bluebook-style cite from either.
edit_highlight answers JSON when called with ?context=viewer and an
HX-Trigger otherwise.
Timeline and witnesses¶
get_facts_data orders by date, time by default and accepts the sort
keys of FactsFilter, through sorting.py: filterset_sort_keys()
derives the legal keys from the filter's order_by field,
stored_sort_key() reads the session key or falls back, and
with_valid_sort() drops a stale key so the rest of the filter still
applies. The same three functions serve highlights, witnesses and saved
cases. Facts link to documents and highlights as sources;
facts/access.py (source_for_fact, label_for_matter) rejects a
source or label posted from another matter. Bulk operations follow the
session selection pattern (toggle_fact_select, select_all_facts over
the visible list, bulk_facts_*). facts_pdf prints exactly what is on
screen: generate_facts_pdf() takes the filtered, sorted list and a
filter_summary from describe_facts_filter() so a partial timeline
says so. Witnesses mirror the pattern and additionally link to highlights
through the apply-witnesses modal.
Labels¶
A label is global (matter null) or the matter's own. _labels_for() in
apps/case/labels/views.py is the one definition of what may go on a
matter's material: the global labels plus that matter's. The apply modal
and add_label_to resolve the posted label id through it, so another
matter's label cannot be attached; remove_label_from resolves through
obj.labels so a leftover label can always be taken off. On a move
between matters (_drop_other_matters_labels(), and the equivalent in
ingest_pdf) the leaving matter's labels stay behind and global labels
travel with the document.
Saved cases¶
caselaws_lookup asks CourtListener for a citation and parks the result
in the session; caselaws_save creates the CaseLaw and runs the summary
task. caselaw_viewer fetches the opinion on demand (fetch_opinion,
falling back to the cluster's first sub-opinion and caching the
opinion_id) and renders html_with_citations with |safe in
templates/case/caselaw-viewer.html. That is a known decision: the
markup is CourtListener's, the viewer anchors highlights on it, and when
the viewers were hardened in October 2026 (highlight JSON and research
excerpts) this one was left as it is pending a real sanitiser. Saved
cases and the viewer need perm_research, enforced by URL pattern in the
middleware.
With AI configured the list is a view of the AI tab:
templates/case/ai/view-pills.html puts a Conversations | Case Law
switch in both toolbars, shown only when caselaw_available (a
CourtListener token, and an admin or a perm_research user).
/case/<id>/caselaws/ and the caselaws tab key render the list with
the AI tab marked active in the case navigation. Without AI the list is
its own Case Law tab, with no AI column, bulk AI menu or summary.
Without a CourtListener token the list, lookup and viewer 404
(CASELAW_PATTERN in PermissionMiddleware). Cases also arrive from the agentic chat's save-caselaw
block (see Agentic chat), which builds the row from
fetch_cluster() so a citation-less slip opinion saves too.
Jurisdictions (apps/case/jurisdictions.py) feed the agent's
search_caselaw court filter: a state lists its supreme and appellate
courts and home circuit, and federal adds the state's district courts,
that circuit and the Supreme Court. District lists are filled in as users
need them (Georgia so far); verify any new id against the CourtListener
courts endpoint, because a typo silently narrows every search that uses
the filter.
Search¶
Search inside a matter (apps/case/search/views.py) is hybrid:
watson.search(query) over the whole index (the case models are
registered in apps/case/search_config.py), keeping the document,
highlight, fact and note hits on this matter, then meaning matches from
the pgvector index (apps/case/ai/semantic.py, semantic_entries())
appended and marked semantic (none without a Gemini key, which the
embeddings need). Highlights on case law are left out
because the result row renders document highlights only. The importance
filter means "at least this important", as on the other tabs. The sidebar
search (apps/search/)
runs watson over matters, contacts, intakes and notes with synonym
expansion, matches proceedings by normalised case number, and shows only
matters the user may open; apps/search/api.py exposes the same search to
the MCP server.
Background work¶
| Task | Queued by | On failure |
|---|---|---|
apps.case.documents.tasks.process_document_ocr |
upload, file replacement, mbox import, Drive ingest, retry and force | ocr_status="failed", ocr_error set, badge offers retry |
apps.case.documents.tasks.generate_document_summary |
the OCR task, only with AI configured | logged, document keeps no summary |
apps.case.caselaws.tasks._generate_caselaw_summary |
caselaws_save and the agent's save-caselaw block, through generate_caselaw_summary(), only with AI configured |
logged, the case keeps no summary |
apps.drive.records.resync_mapping_by_id |
the Drive Folder modal | logged; the nightly full pass catches up |
apps.drive.google.scheduled_sync, scheduled_sync_full |
schedules drive-sync and drive-sync-nightly-full |
a 410 on the cursor re-bootstraps; other errors raise and the next tick retries |
The post_save signal in apps/case/documents/signals.py also queues OCR
on a document created with its file in the first save. Schedules are
listed in the schedules reference.
Access¶
Everything under /case/ passes PermissionMiddleware.process_view,
which resolves every id in the URL (matter, document, highlight, fact,
witness, label, case law, and the object_type/object_id pairs of the
label and witness modals) to a matter through MATTER_LOOKUPS in
apps/accounts/access.py and refuses a user who is not a member of all of
them. The tab views therefore carry only @login_required (the shell
views in views.py that take a matter id also carry
@matter_access_required). What the URL
cannot see is checked in the app: a matter in a form or POST body
(documents/access.py), a source or label id in a body
(facts/access.py, labels/views.py), and lists that would reach across
matters (documents_for_user() on the duplicate warning). Saved cases
and the case-law viewer need perm_research. The
matrix is in the permissions reference.
Things that bite¶
- A new kind of id in a
/case/route needs a line inMATTER_LOOKUPS. Without it the middleware has nothing to check and the view is open to any signed-in user. Document.save()is the only place the proceeding rule runs.queryset.update()skips it (andfull_clean()and history). Write category and proceeding together, asapply_mapping_fields()does.- Fingerprints describe the uploaded bytes, not the stored file. OCR
replaces the file with the OCR'd PDF but leaves
content_hashalone, so a re-upload of the original still matches. A path that replaces bytes on purpose must clear the hash or callset_fingerprints(). - The OCR signal never sees a two-phase upload. The first save has no
file, so the signal marks it
not_applicable; the view must queue OCR itself after the file save, asdocuments_addandrecords._queue_ocr()do. - The mirror never deletes and never writes to Drive. Do not add a
delete path to the sync; delete in the application, which tombstones.
The
drive_file_idunique constraint is what makes the upsert safe under two concurrent passes (IntegrityErroris caught). watson.search()in the matter search is not scoped. It searches the whole index and the view filters by matter afterwards; a big index makes every matter search cost the same.- Sort keys from the session reach
order_by(). A stored key the list does not know would 500 every load; usesorting.pyfor any new sortable list. - The CourtListener throttle is per process.
courtlistener_throttle.pyspaces requests within one process; qcluster workers, the vetting thread pool and the agent's research tools each keep their own clock, so the account's per-minute limit is the real ceiling. - The case-law viewer trusts CourtListener's HTML. Any change to how opinions are fetched must keep that in mind, or add a sanitiser with an allow-list of CourtListener's own elements.
Related¶
- Guide: Documents, Timeline, witnesses and highlights, Case law
- Admin: Google Workspace (the Drive connection), Storage
- Subsystems: Matters, AI context, Agentic chat, Notes and drafts, Email and intakes (the Emails tab), MCP (the case JSON API), Operations
- Conventions: HTMX and Alpine, Session state