How to evaluate statement extraction tools for an RIA
Twelve questions worth asking any vendor in this category, including us. They are written to be answerable in a demo rather than in a brochure, and none of them can be settled with an accuracy percentage.
No products are named on this page, ours included. If you want to apply these to TieoutIQ, the uncomfortable half of the answers is on what TieoutIQ does not do.
The questions
How does the tool prove a figure is right?
Extraction accuracy is usually quoted as a percentage, and a percentage tells you nothing about the statement in front of you. What matters is whether there is an independent reference the extraction is checked against, and the only reference that travels with every statement is the total the statement itself prints.
Ask: What figure do you compare the extraction against, and where does that figure come from?
Is verification a gate or a report?
There is a large difference between a tool that checks its work before handing it to you and one that hands you the numbers and a confidence score. The first cannot ship an unreconciled position set. The second makes that your job.
Ask: If the extraction does not reconcile, can the output still be produced?
What happens to a row it cannot verify?
Silent failure is the actual risk in this category, not a wrong number you can see. A figure that was guessed and presented with the same confidence as one that was checked is the one that ends up in a client document.
Ask: Show me a statement that fails. What does the screen look like, and what can I still export?
Can a figure be traced back to the page it came from?
A compliance spot check is only cheap if the path from a number in the proposal back to the row on the statement is short. Ask separately about computed figures such as fee drag or tax estimates, which derive from the positions rather than appearing on the statement, and therefore trace to inputs rather than to a row.
Ask: Which figures trace to a statement row, and which are computed from them?
What happens to a scan with no text layer?
Plenty of statements arrive as images. A tool that only parses a text layer will fail on them, and a tool that sends everything to a model treats a clean text layer and a photocopy as the same problem. Either is fine if you know which you are buying.
Ask: How is a scanned statement read, and can I tell afterwards which path was used?
How are several currencies handled?
If figures are converted before they are verified, the check is against a number the custodian never published. The order matters, and so does whether the rate, its source and its date survive into the export.
Ask: Is each statement reconciled in its own currency, and is the conversion disclosed per row?
Where is client data processed, and does it train a model?
This is the question that decides whether your CCO signs anything. Get the processing region, the subprocessor list, and the training position in writing, and check whether any part of the pipeline calls a third-party model provider.
Ask: Which regions and which subprocessors? Is there a written commitment that our data trains nothing?
Is there a data processing agreement, and what does the audit trail record?
A DPA you can actually get before a pilot is a reasonable bar. On the audit trail, the useful detail is whether corrections are recorded with the original value and who made the change, not just that an export happened.
Ask: Can I see the DPA now? What exactly is written to the log when someone corrects a row?
Must an adviser review before a client sees the output?
Anything that can generate a client-facing document straight from a machine read is a supervision problem regardless of how good the read is. Review being enforced rather than encouraged is a product decision you can check.
Ask: Can output reach a client without a person signing off on the flagged rows?
Is pricing published, and what does the total cost look like at your size?
Per-account, per-statement and per-seat pricing behave very differently as a book grows. Published pricing also tells you something about how the vendor expects to sell, which affects how much of your time the purchase will take.
Ask: What do we pay in year one at our headcount, including setup, and what changes it?
What does getting data out look like?
Integrations are the usual question, but the more important one is the fallback. If there is no integration with your stack, a complete structured export is the difference between a workable tool and a dead end.
Ask: With no integration to our systems, what file do we get, and does it contain everything?
What happens at the end of a pilot?
Deletion on request, retention defaults, and what happens to uploaded statements if you walk away are all easier to agree before a pilot than after one.
Ask: If we stop, what is deleted, when, and can we get confirmation?
Running the evaluation
Bring a statement that breaks things
A clean quarterly statement from a large custodian tells you very little. Bring the awkward one: a scan, an activity-only document, a household across currencies, or the statement that has cost you an afternoon before. A demo on easy input is a demo of nothing.
Ask to see a failure, not a success
How a tool behaves when it cannot do the job is more informative than how it behaves when it can. Anyone unwilling to show you that is telling you something.
Settle the compliance questions first
Processing region, training commitment, DPA and the supervision path are the items most likely to kill a purchase late. Front-load them and you will not waste a pilot.