Skip to content

Science1 publisher3 min readPublished

UC Berkeley turns 7 million pages of local ordinances into one free searchable database

UC Berkeley researchers turned roughly 7 million pages of local law from all 50 states into a free, searchable database. Comparing codes across towns becomes a search, run on text that an AI pipeline read and sorted.

The Scientist · Science desk

Photograph accompanying UC Berkeley turns 7 million pages of local ordinances into one free searchable database
Photo: berkeley.edu

What happened

  • The raw archive was nearly 10,000 documents, many of them blurry, poorly structured or otherwise inaccessible PDFs.
  • OpenAI models tagged and organized samples of the laws, and those categories were then applied across the whole corpus.
  • Lead researcher Diag Davenport began the work six years ago to measure whether local laws differ in places with high historical discrimination.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • capability A building or nuisance code can be compared across many towns by querying one corpus, and third parties can build local chatbots and custom databases on top of it.
  • constraint Every cross-town comparison inherits the corpus's limit to digitally accessible laws, so places with offline codes drop out of the sample.
  • decision Researchers using the category tags for bias or policy studies will have to run their own accuracy checks before treating tag counts as measurements.
  • cost The heavy OCR compute has already been paid, so downstream users search standardized text at lower cost and never have to re-read 7 million pages of PDFs.

The archive spans all 50 states, but comparing towns depends on how many jurisdictions it covers. By the project's tally, more than 3,000 counties and roughly 6,000 other local governments write their own laws [5], somewhat over 9,000 lawmaking bodies in all [1]. The raw archive was nearly 10,000 documents [4]. Those totals sit close together, but documents and jurisdictions do not map one to one, and the collection takes in only laws that were digitally accessible [1].

The pipeline itself is careful work. Davenport worked with postdoctoral scholar Denis Peskoff, AI researcher Joe Barrow and undergraduate Christopher Vu [17]. The team gathered municipal and county codes from thousands of government and third-party websites, consulting lawyers and designing the collection to meet the hosting sites' technical requirements [6]. The documents averaged about 700 pages each [2], and many were blurry or poorly structured PDFs [4]. LightOnOCR, a vision-language OCR model, extracted both the text and its structure [7]. That step took a large amount of computing power. It produced standardized machine-readable text that is cheaper to search, analyze and hand to AI systems than the original files [8].

Sorting came next. OpenAI models tagged and organized samples of the laws, and the team distilled that work into categories applied across the full corpus [9]. The expensive models read only the samples. The phys.org account does not report error rates for the OCR or the tagging, or say how often the corpus will be refreshed as councils amend their codes.

The question that started the project depends on that precision. Davenport is an assistant professor of technology policy, governance and society in Berkeley's Goldman School of Public Policy and School of Information [10]. He began the project six years ago while studying algorithms, systemic bias and the criminal legal system [11]. He wanted to measure how local laws differ between places with high levels of historical discrimination and places without [11]. No database of all local laws existed. They were scattered across proprietary company databases and convoluted local government websites [12].

"Once you realize how fragmented it all is, it's easy to understand why no one's done the work," Davenport said [13]. The project picked up about a year ago, when advances let the team run one AI tool on top of another [16].

According to phys.org, this is the first database built specifically for local laws; state and federal statutes have had online repositories for years [2]. The account says the resource lets others build custom databases and local chatbots that compare rules on topics such as home construction and dog leashes [15]. Davenport called it "a new kind of telescope" [14]. A researcher running one query across a fixed corpus and a resident asking a chatbot about this year's leash rule need different things from it: the resident also needs the text to be current.

What to watch

  • A published accuracy check of the LightOnOCR extraction and the OpenAI-derived tags against hand-read ordinances.
  • How many of the 9,000-plus lawmaking jurisdictions the corpus actually covers, and how often it is refreshed as codes are amended.
  • First results from Davenport's comparison of local laws in places with high and low historical discrimination.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories