The newsfeed has reached a fever pitch about the dangers of artificial intelligence and “AI slop” in legal practice, with most of it focused on citations or on ensuring that authorities support the proposition advanced. But the OpenAI Agent breach of Medicare shows this is a very narrow view of the promise and peril of AI technologies for the legal profession, writes Chantal McNaught.
At the same time, encouragement to adopt AI wisely, with professional judgement at the forefront, is coming from both the Federal Court of Australia and the Fair Work Commission.
The real risks
Hyperfocusing on drafting errors creates a false sense of security that masks the deeper technological hazards for the profession, making it more frustrating for small and medium-sized firms to adopt the technology responsibly.
A lawyer’s foundational duty is to the administration of justice. This is an absolute duty which cannot be outsourced. Taking vendor assurances at face value for increasingly complex AI capabilities without interrogating those same system mechanics would be an abdication of the duty to the administration of justice. This is why developing both literacy (understanding what the technology can do) and fluency (understanding how to use it effectively, efficiently, ethically, and safely) is so critical.
Large Language Models (LLMs) and Generative Pre-Trained Transformers (GPTs) of the Claude and ChatGPT variety are known for drafting text after being “prompted” by a human to do so. Agentic systems are a level of complexity higher than this, which allow AI agents to pursue objectives independently – hence “agents”.
Agentic AI’s risk profile is a fundamental shift away from generating legal text to executing unmonitored actions. But the answer is not to reject AI or embrace it without management.
What we know so far of the Medicare statistics portal breach and DSEwiki disclosures is that goal-oriented AI agents can and do actively bypass external security controls to retrieve restricted data to achieve their objectives. Agents on a simple research mission conducted by OpenAI’s labs were looking for data to support their objectives. This is what unsupervised, unharnessed and unguarded machine conduct looks like. It can engage in harmful security practices without human intent.
This kind of agentic risk renders the “human-in-the-loop” a cliché. Reviewing final drafts prior to court submission only catches the errors on the page (and manages those risks). Such a post-execution human checkpoint does not mitigate the risk when an AI agent has breached networks, accessed private logs, or collected unauthorised data to complete its objectives of generating the requested text. There is often little transparency when an AI agent is tasked with finding information; these recent events highlight that the agent can execute internet searches and target systems to retrieve information it has been tasked to retrieve.
Why small firms are exposed
These are not hypotheticals, and I’m most concerned about the small- to medium-sized law firms that employ roughly 94% of all Australian lawyers. The technology stack (the digital tools and products a firm uses to provide legal services) has grown more complex over the past 5 years alone. Add to that the deluge of business administration presented by changes to Privacy Act obligations, Anti-Money Laundering requirements, and director obligations, and the risk exposure rises considerably when firms deploy agentic workflows at vendors' say-so; the burden on these small- to medium-sized firms is exhausting.
There is also a real temptation to deploy systems that promise fast, cheap, high-quality outputs to enhance legal services. There is a very real, underserved market for legal services, and lawyers tell me they are under more pressure than ever to provide the trifecta of fast, cheap, and high-quality legal services. This temptation to finally provide all three to clients can feel like providing true value for the administration of justice. But this promise comes with peril.
What to ask vendors
Accepting a vendor's euphemism for “misaligned model activity” from an internal investigation directly compromises the ethical obligations of Legal Practitioner Directors at a time when they are already stretched. When a vendor is pushing the benefits of agentic AI in legal practice, it is prudent to ask about the types of transparency, harnesses, and limitations the vendor has already put in place on top of frontier AI model technologies. Does this tool support internet search and agentic workflows? What limitations can I add to the tool to handle a sensitive task? Are there action logs, and can I access them? These questions can help not just with procurement decisions but also with applications and risk management for the firm.
Lawyers must not succumb to the make-believe dichotomy between AI hype and AI doom. Standard AI use for text generation is managed well with governance small firms can manage. The new risks of agentic AI in legal practice present an opportunity for the profession to demand transparency from vendors deploying agentic technologies targeted at legal practice. Demand execution logs, enforce hard operational boundaries, consider lobbying for regulation if necessary, and refuse to deploy autonomous agents until a requisite level of responsibility can be enforced. To do so requires lawyers and leaders to obtain sufficient levels of AI literacy and fluency.
Without adequate fluency, lawyers (and firms) risk gambling the foundational duty to the administration of justice. But building AI fluency in legal practice is attainable and should not be intimidating for small- to medium-sized firms. Many conferences across the nation are now including AI literacy or fluency content in their programming as standard. Law Societies, such as the Queensland Law Society, also provide AI literacy and fluency resources for members. Reflection is also a highly valuable professional attribute for building literacy and fluency skills. This can be as simple as documenting three questions every time AI is used in legal work: What was the outcome of using AI in this task? What did I expect from my use of AI in this task? Was there anything that surprised me or that I would do differently next time?
Chantal McNaught is a PhD candidate at Bond University and director of advisory at 43° Below.