AI & innovation September 26, 2025

How We Developed an AI for Automated Document Validation

Behind the Scenes: Developing an AI for Automated Document and Supporting Evidence Validation

How We Developed an AI for Automated Document Validation

From Initial Failure to Production Success: A Real-World R&D Case Study

Picture this: your customers need to submit supporting documents to access an additional service. Invoices, certificates, attestations, these documents arrive in every possible format: scanned PDFs, smartphone photos, screenshots. Each one contains critical details to verify: the applicant’s name, date, company, stamp, signature, product codes. The challenge? Automating this validation process, which used to take your teams hours, while maintaining flawless accuracy. The catch? Our first AI automation attempt didn’t deliver the expected results. Instead of giving up, we launched a structured R&D project to understand why, and how to improve our approach.

Our AI Document Validation R&D Project in 6 Steps

Step 1: Establishing the Ground Truth

Before we could measure our AI’s performance, we needed to define what a "correct answer" looked like. That’s the principle of ground truth, literally the baseline reality. Here’s how we did it: a human expert manually analyzed around thirty representative documents, meticulously documenting for each one: ✅ Is the applicant’s name present? Missing? ✅ Is the document date visible? Illegible? ✅ Is the company mentioned? Missing? ✅ Is the official stamp present? Missing? ✅ Is the signature identifiable? Not detected? ✅ Is the product code correct? Incorrect?

This step may seem tedious, but it’s absolutely critical. Without this reliable reference, there’s no way to tell if our AI is improving or regressing!

Step 2: Designing and Configuring Different Approaches

Rather than betting on a single solution, we developed two distinct analysis pipelines, each capable of combining multiple specialized AI models. We manually tested each pipeline on our reference documents, experimenting with different parameters and model combinations to identify the most promising configurations.

Step 3: Running Large-Scale Tests

Once our pipelines were configured, we ran them across all 30 reference documents. Every result was automatically stored in a database for comparative analysis.

Step 4: Comparing Against Reality

Now comes the moment of truth: matching our AI’s predictions against the ground truth established by the human expert. For each document and every piece of information we were looking for, we now know whether our system got it right, or wrong.

Step 5: Evaluating Performance with Business Metrics

To objectively assess our results, we use four standard AI performance indicators. Here’s what they mean in practical terms:

Accuracy

“Out of 100 checks, how many are correct?” If our AI examines 100 elements and gets 85 right, the accuracy is 85%. It’s the most intuitive metric, but not always the most relevant.

Precision

“When the AI says ‘present,’ how often is it right?” If the AI detects 20 stamps and 18 are actually present, the precision is 90%. This metric measures the reliability of positive detections.

Recall

“Out of all the elements that are actually present, how many does the AI find?” If 25 documents actually contain a stamp and the AI detects 20, the recall is 80%. This metric measures the ability to miss nothing important.

F1-Score

“What’s the overall balance between precision and recall?” The F1-score combines precision and recall into a single metric. The closer it is to 100%, the better. It’s often the go-to indicator for comparing different approaches.

Step 6: Comparing and Analyzing Results

We calculated these metrics from two perspectives:

Overall Pipeline Performance

AI document validation

Which pipeline delivers the best overall performance? Pipeline A or Pipeline B?

Detailed Breakdown by Information Type

AI document validation

For each piece of information we’re looking for (name, date, stamp, etc.), which pipeline is the most reliable? What are our weak spots to monitor? This dual analysis allowed us to choose the best configuration and pinpoint areas requiring extra attention in production.

Mission Accomplished!

Thanks to this methodical approach, we successfully built and validated a pipeline robust enough for production deployment. Our system now automates the first step of document validation, with a compliance level that lets us handle most cases independently. That said, we’re keeping a close eye on one or two specific aspects identified during our analysis, ensuring optimal service quality. The result? Faster processing for your customers, fewer repetitive tasks for your teams, and reliability backed by science.

Facing AI Challenges in Your Project?

This methodical approach, from problem analysis to validation in production, shows how we design artificial intelligence to enhance your customer relationships. Got an idea? A project? Let’s talk! At fAibrik, we turn your technical challenges into concrete solutions, combining scientific rigor with the simplicity of a service that just works.

Want to go further?