{"id":13792,"date":"2023-02-18T17:19:48","date_gmt":"2023-02-18T17:19:48","guid":{"rendered":"https:\/\/prizmlaw.com\/site\/?p=13792"},"modified":"2025-04-18T17:52:26","modified_gmt":"2025-04-18T17:52:26","slug":"dev-of-lang-models","status":"publish","type":"post","link":"https:\/\/prizmlaw.com\/site\/2023\/02\/18\/dev-of-lang-models\/","title":{"rendered":"Development of Language Models to Process Legal Language"},"content":{"rendered":"\t\t<div data-elementor-type=\"wp-post\" data-elementor-id=\"13792\" class=\"elementor elementor-13792\">\n\t\t\t\t<div class=\"elementor-element elementor-element-bab31eb e-flex e-con-boxed e-con e-parent\" data-id=\"bab31eb\" data-element_type=\"container\">\n\t\t\t\t\t<div class=\"e-con-inner\">\n\t\t\t\t<div class=\"elementor-element elementor-element-2b61188 elementor-widget elementor-widget-pix-img\" data-id=\"2b61188\" data-element_type=\"widget\" data-widget_type=\"pix-img.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t<div class=\"pix-img-element d-inline-block \" ><div class=\"pix-img-el    center d-inline-block  w-100 rounded-lg\"  ><img fetchpriority=\"high\" decoding=\"async\" class=\"card-img2 pix-img-elem rounded-lg  h-1002\" style=\"height:auto;\" width=\"990\" height=\"632\" srcset=\"https:\/\/prizmlaw.com\/site\/wp-content\/uploads\/2025\/04\/robotJudge.jpg 990w, https:\/\/prizmlaw.com\/site\/wp-content\/uploads\/2025\/04\/robotJudge-300x192.jpg 300w, https:\/\/prizmlaw.com\/site\/wp-content\/uploads\/2025\/04\/robotJudge-768x490.jpg 768w\" sizes=\"(max-width: 990px) 100vw, 990px\" src=\"https:\/\/prizmlaw.com\/site\/wp-content\/uploads\/2025\/04\/robotJudge.jpg\" alt=\"Image link\" \/><\/div><\/div>\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t<div class=\"elementor-element elementor-element-0ca989f elementor-widget elementor-widget-heading\" data-id=\"0ca989f\" data-element_type=\"widget\" data-widget_type=\"heading.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t<h2 class=\"elementor-heading-title elementor-size-default\">Development of Language Models to Process Legal Language<\/h2>\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t<div class=\"elementor-element elementor-element-55fee98 e-flex e-con-boxed e-con e-parent\" data-id=\"55fee98\" data-element_type=\"container\">\n\t\t\t\t\t<div class=\"e-con-inner\">\n\t\t\t\t<div class=\"elementor-element elementor-element-c5caf4c elementor-widget elementor-widget-text-editor\" data-id=\"c5caf4c\" data-element_type=\"widget\" data-widget_type=\"text-editor.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t\t\t\t\t<p>One possibility that has intrigued me for years at the intersection of law and technology has been getting computer systems to \u201cunderstand\u201d the substance of all kinds of contracts. This could open up a world of possibilities for automating legal processes and making legal services more accessible.<\/p><p>Imagine being able to ask your virtual assistant, \u201cHey Siri\/Alexa, I\u2019m thinking of moving to Chicago. What contracts that I\u2019m party to would prevent this or would need to be updated?\u201d The system would be aware of all your personal agreements and provide a synthesized answer, including your employment contract, company policies, car lease, apartment lease, health, car, life, renter\u2019s insurance, and service agreements, including cell phone, internet, cable, etc. The system could even make the necessary adjustments for you, including giving appropriate notice to withdrawal from some agreements and entering into new ones, such as a new lease on an apartment. This idea is not as far-fetched as it was just a few years ago.<\/p><p>At the core of this idea is the challenge of getting computer systems to understand the substance of contracts. Most contracts exist as unstructured data to a computer system. Up until about five years ago, there were two main ways to go about building a system that could understand contracts.<\/p><p>The first approach involves developing a law-specific programming language. Since much of a contract can be broken down into if\/then statements, then perhaps instead of (or in addition to) writing contracts in English, contracts could also be expressed in a type of programming language, then computers could understand contracts. These already exist in finance and in some insurance settings. However, this approach depends on the contract being originally expressed in a machine-readable format at the time of drafting. The <a href=\"https:\/\/law.stanford.edu\/codex-the-stanford-center-for-legal-informatics\/\" target=\"_blank\" rel=\"noreferrer noopener\" data-type=\"URL\" data-id=\"https:\/\/law.stanford.edu\/codex-the-stanford-center-for-legal-informatics\/\">Codex<\/a> center at Stanford has some projects related to this generally referred to as <a href=\"https:\/\/youtu.be\/PsSAacByrB0\" target=\"_blank\" rel=\"noreferrer noopener\" data-type=\"URL\" data-id=\"https:\/\/youtu.be\/PsSAacByrB0\">computable contracts<\/a>.<\/p><p>The second approach involves using machine learning to gather a bunch of contracts, break them down into the important terms and clauses, label all of that data, and build an ML model based on that data. That contract model could then be used to analyze new contracts. <a href=\"https:\/\/www.lawgeex.com\/\" data-type=\"URL\" data-id=\"https:\/\/www.lawgeex.com\">LawGeex<\/a> is an example of a company that has been using this approach.<\/p><p>While the machine learning approach is more doable than creating a law-specific programming language, it still requires a ton of work gathering and labeling legal language, just to build what might be a fragile AI model. There are also challenges in terms of understanding the nuance and context of legal language, as well as the potential for bias in the data used to train the AI model.<\/p>\t\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t<div class=\"elementor-element elementor-element-8e87c26 elementor-widget elementor-widget-heading\" data-id=\"8e87c26\" data-element_type=\"widget\" data-widget_type=\"heading.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t<h2 class=\"elementor-heading-title elementor-size-default\">Evolutions in Natural Language Processing<\/h2>\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t<div class=\"elementor-element elementor-element-04d1695 elementor-widget elementor-widget-text-editor\" data-id=\"04d1695\" data-element_type=\"widget\" data-widget_type=\"text-editor.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t\t\t\t\t<p>All of this takes place in the context of what is known in the AI world as natural language processing (NLP). One of the earliest breakthroughs in NLP was the release of <a href=\"https:\/\/towardsdatascience.com\/word2vec-explained-49c52b4ccb71\" target=\"_blank\" rel=\"noreferrer noopener\" data-type=\"URL\" data-id=\"https:\/\/towardsdatascience.com\/word2vec-explained-49c52b4ccb71\">word2vec<\/a> in 2015. This learning algorithm helped computers understand semantic relationships between words by learning vector representations of words.<\/p><p>Word vectors are important in NLP because they provide a way for computers to understand the meaning and context of words. By representing words as vectors, NLP models can perform mathematical operations on them, such as addition and subtraction, to infer relationships between words. For example, the vector for \u201cking\u201d minus the vector for \u201cman\u201d plus the vector for \u201cwoman\u201d would result in a vector close to the vector for \u201cqueen\u201d. This ability to capture semantic relationships between words was a significant advancement in the field of NLP and opened up new possibilities for applications such as text classification and information retrieval.<\/p><p>Additionally, transfer learning became a dominant paradigm in NLP with the release of <a href=\"https:\/\/towardsdatascience.com\/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270\" target=\"_blank\" rel=\"noreferrer noopener\" data-type=\"URL\" data-id=\"https:\/\/towardsdatascience.com\/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270\">BERT<\/a> in 2019, which pre-trains a large neural network on a vast corpus of text data and fine-tunes it for a specific downstream task. More recent advancements include the release of <a href=\"https:\/\/openai.com\/api\/\" target=\"_blank\" rel=\"noreferrer noopener\" data-type=\"URL\" data-id=\"https:\/\/openai.com\/api\/\">GPT-3<\/a> in 2020, which is a neural language model with 175 billion parameters (these large language models are often referred to as LLMs). It achieved impressive results on various language tasks such as language translation and text completion. This model became super-popular in 2021 with the release of <a href=\"https:\/\/chat.openai.com\/\" target=\"_blank\" rel=\"noreferrer noopener\" data-type=\"URL\" data-id=\"https:\/\/chat.openai.com\">ChatGPT<\/a> \u2013 a public application based on an underlying GPT model.<\/p><p>With this as a background, <a href=\"https:\/\/prizmlaw.com\/site\/2025\/04\/18\/using-lang-models-1\/\" data-type=\"post\" data-id=\"2049\">in my next article<\/a> I\u2019ll walk through using GPT-3 and Python to break down complex legal documents for basic question and answering.<\/p>\t\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t","protected":false},"excerpt":{"rendered":"<p>Development of Language Models to Process Legal Language One possibility that has intrigued me for years at the intersection of law and technology has been getting computer systems to \u201cunderstand\u201d the substance of all kinds of contracts. This could open&#8230;<\/p>\n","protected":false},"author":1,"featured_media":13794,"comment_status":"closed","ping_status":"open","sticky":false,"template":"elementor_header_footer","format":"standard","meta":{"_siteseo_robots_primary_cat":"4","pagelayer_contact_templates":[],"_pagelayer_content":"","footnotes":""},"categories":[27],"tags":[],"class_list":["post-13792","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-legal-tech"],"_links":{"self":[{"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/posts\/13792","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/comments?post=13792"}],"version-history":[{"count":7,"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/posts\/13792\/revisions"}],"predecessor-version":[{"id":13821,"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/posts\/13792\/revisions\/13821"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/media\/13794"}],"wp:attachment":[{"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/media?parent=13792"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/categories?post=13792"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/prizmlaw.com\/site\/wp-json\/wp\/v2\/tags?post=13792"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}