Social Sciences › Social Sciences › Political Science and International Relations
Artificial Intelligence in Law
271 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States34% · 47 papers
- China16% · 22 papers
- Germany13% · 18 papers
- India12% · 16 papers
- France5.1% · 7 papers
- United Kingdom5.1% · 7 papers
- Canada4.4% · 6 papers
- Japan4.4% · 6 papers
Across 137 papers on this subject with at least one lab located. 44 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Law And Order: Tax Law Autoformalization
Sophia Simeng Han, Yoshiki Takashima, Anjiang Wei, Zhaoyu Li, Michael Genesereth · 5 October 2026
Legal systems are increasingly implemented through software, yet scalable methods for translating legal texts into accurate symbolic representations remain underdeveloped. We study this problem through tax law, where forms and filing instructions define large computational structures involving arith…
- Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System
Rub\'en Manrique, Michelle Castellanos, Jorge Morales, Juan David Guti\'errez, Antonio Barreto Rozo, Joaqu\'in V\'elez Navarro · 5 October 2026
Large language models (LLMs) are increasingly used to support legal practice, education, and research, yet their reliability in national legal systems outside the United States remains largely undocumented. We introduce an expert-validated benchmark for evaluating LLM reliability on the Colombian le…
- ARCCS: An Automated Regulatory Compliance Checking System
Giorgos Filandrianos, Jos\'e Menezes, Chrysoula Zerva, Alessandro Gianola · 2 October 2026
Regulatory compliance checking - deciding whether a target document satisfies the obligations of a regulation - requires interpreting dense legal text, identifying which provisions apply, and grounding each decision in explicit evidence. We present ARCCS, an end-to-end, automated, agentic, and regul…
- LawCompass: Navigating from Legal QA to Multi-Agent Deep Research with Grounded Evidence
Xiaoxia Cheng, Linnan Wang, Jiahao Ma, Zhichuan Ye, Xuemei Zhou, Chuanyu Tong, Bo Jiang, Qing Zhu · 2 October 2026
Recent advances in Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have significantly democratized access to legal information. Nevertheless, most existing legal assistants remain confined to multi-turn conversational QA, failing to support complex legal tasks that require syst…
- Legal text classification in Korean sexual offense cases: from traditional machine learning to large language models with XAI insights
Jeongmin Lee · 2 October 2026
The advancement of natural language processing (NLP) has expanded AI-based text classification in the legal domain. However, accurately classifying legal documents remains challenging due to the complexity of legal texts and subtle differences between legal categories. This study evaluates legal tex…
- The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating
Gnaneswar Villuri, Hashmath Shaik, Alex Doboli · 2 October 2026
Reading an LLM judge's verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihood-scoring evaluation harnesses produce. We show that this readout distorts position bias in one direction: it overstates it in every con…
- JusticeAxis: Benchmarking Legal Judgment between Rigid Rule Application and Ungrounded Discretion
Zhengkai Tu, Mingda Zhang, Zijia Wang, Xiaoying Tang, Jimmy Huang · 2 October 2026
A sound judgment applies the law to established facts and weighs the circumstances in which they arose. However, existing methods swing between rigid statute matching and ungrounded discretion, benchmarks score a label or a rubric, and the experience that would supply the balance stays unverified. W…
- When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts
Huichan Seo · 2 October 2026
Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 cult…
- Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents
Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan · 2 October 2026
Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even pa…
- LAURA: Knowledge Distillation for Interpretable Ambiguous Clause Identification in Legal Contracts
Amrita Singh, Aditya Joshi, Jiaojiao Jiang, Hye-young Paik · 1 October 2026
Legal contracts contain ambiguities that expose enterprises to financial and legal risks. Some ambiguities allow flexible interpretation without triggering disputes, while others lead to significant legal conflicts. This makes identification alone insufficient, and interpretable rationale analysis e…
- No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
Michael Krumdick, Charles Lovering, Varshini Reddy, Seth Ebner, Chris Tanner · 30 September 2026
Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is appealing due to its scalability, low cost, a…
- RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications
Jianpeng Zhao, Haihua Xu, Haoyang Zhang, Shuang Qian, Yixiang Tang, Xintao Wang, Kun Sun, Pei Wu, Shuhan Zhong, Pengyang Wang · 30 September 2026
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require …
- SinBrief: A Hybrid Framework for Abstractive Text Summarisation of Sinhala Legal Documents
Minduli Lasandi, Nevidu Jayatilleke · 30 September 2026
Legal document summarisation in low-resource languages presents significant challenges due to the scarcity of annotated data and the complexity of domain-specific terminology. This paper presents SinBrief, a hybrid abstractive summarisation framework for Sinhala legal documents that does not require…
- JudgeProfile: Understanding and Steering Subjectivity in LLM Judges
Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He · 30 September 2026
LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specifi…
- A decision-support system applied to Law: Reasoning and explainability of the decision
Jeremy Bouche-Pillon (IRIT, IRIT-MELODI, IRIT-ADRIA, IRIT-LILaC), Pascale Zarat{\'e} (IRIT, UT Capitole, IRIT-ADRIA), Yannick Chevalier (IRIT-MELODI, IRIT, CNRS), Nathalie Aussenac-Gilles (IRIT-MELODI, IRIT, CNRS) · 29 September 2026
The emergence of the digital transition brought an increasing need to control the processing of digital information, including in Law Enforcement Agencies (LEAs). At the EU level, in recent years, many regulations have emerged to control data processing and exchange. Texts other than the GDPR, such …
- Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict
Sourabrata Mukherjee, Sunayana Sitaram · 29 September 2026
LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal re…
- LLM Judge Validation Under Sparse Overlap: From Inference to Design
Junxuan Li, Arko Mukherjee, Soumyabrata Pal · 29 September 2026
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this \emph{overlap sparsity} is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates re…
- The Death of the Legal Author. Authority, Intention, And Law-Creation in the Advent of GenAI
Julieta A. Rabanos, Bojan Spai\'c · 29 September 2026
Generative artificial intelligence in the form of chatbots based on large language models (LLMs) has taken the world of law by storm. Philosophy of law is struggling to catch up with the theoretical significance of the advent of technological development and the way it may modify traditionally estab…
- Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
Yongfu Zhu, Lin Sun, Jinzhu Wu, Weihong Lin, Xiaoqi Jian, Guangxiang Zhao, Change Jia, Linglin Zhang, Sai-er Hu, Yuhan Wu, Xiangzheng Zhang · 28 September 2026
Reasoning models represented by the Deepseek-R1-Distill series have been widely adopted by the open-source community due to their strong performance in mathematics, science, programming, and other domains. However, our study reveals that their benchmark evaluation results are subject to significant …
- ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints
Sriram Kannan, Swetha Saseendran, Vishnu Vardhan Reddy Kandi, Leslie Barrett, Madhavan Seshadri, Enrico Santus · 25 September 2026
U.S. employment-discrimination complaints describe complex event sequences that are not explicitly captured by lexical or embedding-based representations alone. We present ARGUS, a source-grounded pipeline that combines a 5W1H-inspired schema, legal-domain models, and LLM-based structured generation…
- Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation
Tobias Deu{\ss}er, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa · 25 September 2026
Financial institutions operate under dense, frequently amended rulebooks, and answering a compliance question correctly requires not only fluency but verifiable grounding in the authoritative text. Large language models are attractive for this task, yet the models that firms can realistically deploy…
- ContraVis: Evidence-Grounded Visual Analytics for Contradiction Review in Legal Contracts
Luis Sante, Paula Lima, Mariana Rocha, Jorge Poco · 24 September 2026
Legal contracts are structurally complex documents in which contradictions may emerge across distant and interconnected provisions. Although large language models (LLMs) improve legal language understanding, contradiction analysis remains a human-centered and evidence-grounded review task. We presen…
- LabourCrew: A Multi-Agent RAG Framework for Trustworthy Adversarial Deliberation and Statutory Reasoning over Labour Law
Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain · 24 September 2026
In statutory question answering, every claim must be traceable to evidence, not merely relevant, since unverifiable labour-rights answers carry serious legal consequences. Current systems fall short: single-pass RAG cannot detect insufficient evidence, while multi-agent legal-debate systems treat gr…
- Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding
Fan Zhang, Yankai Chen, Zhuohan Xie, Yixi Zhou, Sijia Peng, Lei Fan, Xinhua Ji, Cunyuan Zheng, Huangyong Shan, Philip S. Yu, Xue Liu, Yu Chen, Preslav Nakov, Songwei He · 24 September 2026
Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on Contra…
- Cross-Lingual Legal QA for Vietnamese Labour Law: Retrieval, Translation, and Verifier-Guided Correction
Nguyen Minh Chi, Mo El-Haj, Nguyen Ha Thanh, Dawn Knight, Paul Rayson · 24 September 2026
Cross-lingual legal question answering must retrieve statutes across languages while preventing unsupported legal claims. We introduce a bilingual evaluation suite of 231 Vietnamese--English question--answer pairs from Vietnamese labour law. Of these, 75 are additionally annotated for five challengi…
