Go Back

Beyond the Legal AI Demo

After testing two specialist legal AI tools, Spellbook and Wordsmith AI over a three week trial period, we came away with plenty of lessons.

How we trialled Spellbook and Wordsmith AI

We’ve been using our own internal bespoke and general-purpose AI tools, but wanted to  assess specialist legal tools to understand the capabilities available on the market. 

Watching demos or attending webinars hosted by these providers only gets you a quick look at their features, and in most instances, it’s a demo on a standard NDA, which doesn't answer questions we have about the level of complexity these tools can engage with. It also doesn't tell you how they integrate with our existing assets (e.g., templates, playbooks, data extraction tools and standard operating procedures).

We decided on testing Spellbook and Wordsmith AI, because we’re looking for tools focused on contract review and drafting, as opposed to CLM tools. We have our own existing technology that we’ve used and refined over the years that tracks and follows the contracting lifecycle. 

 I’ve shared some notes on the process and outcome below. 

Scope of testing

We decided on three testing groups: sales, procurement and non-disclosure agreements. This gave us an opportunity to test the agreements where we see high volumes, while at the same time addressing more complex and heavily negotiated agreements. We also decided to have dedicated teams test each agreement type but implemented standardised feedback sheets, so the results could be compared across the work streams. 

Testing window

We ran both tools in parallel. Only one stream tested on live matters, in partnership with a client; the rest worked from agreements prepared in advance. It took careful planning to make the most of the testing time available, with regular checkpoints to shift focus if necessary.. 

Benchmarking

It was clear to us from the outset that we needed to compare the results against our internal AI tool, Ray, and CoPilot, both of which our whole team uses day-to-day. We also captured our baseline manual review time, and it’s worth noting that we’ve been using technology solutions across workflows to improve productivity long before AI arrived, so we’re measuring against an already optimised process. 

Preparation

We focused on getting four things ready before testing:

  1. Optimised playbooks for issue spotting and redlining;
  2. Optimised templates;
  3. A set of agreements to test on;
  4. A set of agreements to use for data extraction. 

We work with optimised playbooks and templates as standard, so the prep wasn't a heavy lift for us and it turned into a useful exercise in its own right, surfacing improvements to both. 

If you don’t have this in place, it’s a good idea to start here before you kick off a trial. Otherwise, you will be using valuable testing time to figure out how to get your playbook into a format that the tool understands, or scrambling to find good examples of agreements you want to test on. We recommend doing this work upfront so you can get straight into testing when the trial starts.

Findings

Playbook onboarding

While you can use AI to get your playbook into a shape that’s easier for the tools to process, you won’t be able to avoid reviewing the rules to ensure they make sense. It’s probably a good idea to do a first version of the playbook manually in a format that can be uploaded to your tool of choice. That makes it much easier to review the rules once they pull through. Wordsmith offers a feature that can convert any MS Word playbook into the format their tool requires and you can amend the rules it creates by prompting Wordsmith the chat function. It will then direct you to any issues in the logic of the relevant rules. Both tools also have a standard template you can use and build the playbook offline and then upload all your rules at once. 

Getting your playbook rules back out of both platforms in a format your team can actually reuse, is still messy across the board. 

Issue spotting

The tools were fairly evenly matched here. All the tools we tested, including Ray and CoPilot  were able to produce issue lists with low setup time. The tricky part is getting the format of how it’s presented right, and it will depend on the audience. Meeting people where the work is done is key here, because an additional step/workflow will ultimately slow you down, even if it adds value from a quality point of view. 

Redlining

Wordsmith and Spellbook  had great functionality on this front, especially the ability to tweak the rules behind the redlines. Setting up your playbook correctly here is important and the output needs to be refined by updating the rules as you test. It will definitely add value on first pass reviews and make reviewing agreements quicker, but manual review of the output will still be required to ensure consistency and accuracy. 

Template generation

We've been using document automation for years, so this was an interesting one to test. We didn't record any major wins on speed, so the jury's still out.

Data extraction and analysis

No tool we've used before comes close to what AI does here. Across the board, AI gets this done efficiently, and we're glad to stop doing this manually. The data extraction and analysis capability tested in Wordsmith also seems more advanced than what CoPilot has to offer. 

Key takeaways

  • Getting your assets ready to use on AI tools is non-negotiable. Whether you pay for it to be done or take on the task internally, it is the first step and skipping this one will cost you. 
  • Be clear on what assets you can export from any tool you use. Improvements should ideally be made to the asset within the AI tool environment, but it’s important that you’re able to access the improved asset afterwards, otherwise you are back to square one if you decide to switch providers. 
  • Testing on the use cases that you’ve been grappling with internally is the most useful approach. That way, you know what you can do realistically without AI or specialist legal AI tools, and you’re in a good position to assess the value it can add. 
  • Don’t forget to record your baseline. Record how long it takes to do the work manually, time spent on quality reviews, and any improvements on time and quality with general-purpose AI tools. 

The tools aren't cheap, and three weeks of testing wasn't enough to prove a return on investment.  Particularly when we're measuring against workflows we've already spent years making more productive.  It's also worth saying that CoPilot held its own and shouldn't be overlooked.

We think one of the tools warrants a longer, closer look before we commit.  In the meantime, the trials gave us what we actually needed: a documented baseline, a better view of where AI belongs in our workflows, and enough understanding to match the tool to the task rather than the other way round.

Keep the momentum going! Here's what to do next:

You might be wondering what your next steps should be. Let us guide you with three easy options:

Take the Optimised Contracting Assessment
Start Quiz
Attend a Webinar
Register
Get Expert Support
Contact Us
Previous article
There is no previous article.
Next article
There is no next article.