After testing two specialist legal AI tools, Spellbook and Wordsmith AI over a three week trial period, we came away with plenty of lessons.
We’ve been using our own internal bespoke and general-purpose AI tools, but wanted to assess specialist legal tools to understand the capabilities available on the market.
Watching demos or attending webinars hosted by these providers only gets you a quick look at their features, and in most instances, it’s a demo on a standard NDA, which doesn't answer questions we have about the level of complexity these tools can engage with. It also doesn't tell you how they integrate with our existing assets (e.g., templates, playbooks, data extraction tools and standard operating procedures).
We decided on testing Spellbook and Wordsmith AI, because we’re looking for tools focused on contract review and drafting, as opposed to CLM tools. We have our own existing technology that we’ve used and refined over the years that tracks and follows the contracting lifecycle.
I’ve shared some notes on the process and outcome below.
We decided on three testing groups: sales, procurement and non-disclosure agreements. This gave us an opportunity to test the agreements where we see high volumes, while at the same time addressing more complex and heavily negotiated agreements. We also decided to have dedicated teams test each agreement type but implemented standardised feedback sheets, so the results could be compared across the work streams.
We ran both tools in parallel. Only one stream tested on live matters, in partnership with a client; the rest worked from agreements prepared in advance. It took careful planning to make the most of the testing time available, with regular checkpoints to shift focus if necessary..
It was clear to us from the outset that we needed to compare the results against our internal AI tool, Ray, and CoPilot, both of which our whole team uses day-to-day. We also captured our baseline manual review time, and it’s worth noting that we’ve been using technology solutions across workflows to improve productivity long before AI arrived, so we’re measuring against an already optimised process.
We focused on getting four things ready before testing:
We work with optimised playbooks and templates as standard, so the prep wasn't a heavy lift for us and it turned into a useful exercise in its own right, surfacing improvements to both.
If you don’t have this in place, it’s a good idea to start here before you kick off a trial. Otherwise, you will be using valuable testing time to figure out how to get your playbook into a format that the tool understands, or scrambling to find good examples of agreements you want to test on. We recommend doing this work upfront so you can get straight into testing when the trial starts.
While you can use AI to get your playbook into a shape that’s easier for the tools to process, you won’t be able to avoid reviewing the rules to ensure they make sense. It’s probably a good idea to do a first version of the playbook manually in a format that can be uploaded to your tool of choice. That makes it much easier to review the rules once they pull through. Wordsmith offers a feature that can convert any MS Word playbook into the format their tool requires and you can amend the rules it creates by prompting Wordsmith the chat function. It will then direct you to any issues in the logic of the relevant rules. Both tools also have a standard template you can use and build the playbook offline and then upload all your rules at once.
Getting your playbook rules back out of both platforms in a format your team can actually reuse, is still messy across the board.
The tools were fairly evenly matched here. All the tools we tested, including Ray and CoPilot were able to produce issue lists with low setup time. The tricky part is getting the format of how it’s presented right, and it will depend on the audience. Meeting people where the work is done is key here, because an additional step/workflow will ultimately slow you down, even if it adds value from a quality point of view.
Wordsmith and Spellbook had great functionality on this front, especially the ability to tweak the rules behind the redlines. Setting up your playbook correctly here is important and the output needs to be refined by updating the rules as you test. It will definitely add value on first pass reviews and make reviewing agreements quicker, but manual review of the output will still be required to ensure consistency and accuracy.
We've been using document automation for years, so this was an interesting one to test. We didn't record any major wins on speed, so the jury's still out.
No tool we've used before comes close to what AI does here. Across the board, AI gets this done efficiently, and we're glad to stop doing this manually. The data extraction and analysis capability tested in Wordsmith also seems more advanced than what CoPilot has to offer.
Key takeaways
The tools aren't cheap, and three weeks of testing wasn't enough to prove a return on investment. Particularly when we're measuring against workflows we've already spent years making more productive. It's also worth saying that CoPilot held its own and shouldn't be overlooked.
We think one of the tools warrants a longer, closer look before we commit. In the meantime, the trials gave us what we actually needed: a documented baseline, a better view of where AI belongs in our workflows, and enough understanding to match the tool to the task rather than the other way round.
You might be wondering what your next steps should be. Let us guide you with three easy options: