n 3,480 human evaluations
Holistic extraction accuracy across 140 supplier documents.
n 140 documents · four weeks
Analyst time recovered over the pilot window.
145.9 s average · n 140
Processing time per document, ingestion to record.
against $13.44 operating cost
Avoided processing cost over the pilot, estimated on the stated assumption — not cash recovered. Platform operating cost: $13.44; implementation and deployment investment separate.
One trusted record, assembled by hand.
RCMA Makeup qualifies every raw material against the documents its suppliers provide. Ingredient intelligence is the work of turning those documents into one record the business can rely on.
Certificates of analysis, technical and safety data sheets, supplier regulatory statements — each carrying part of the picture, across 29 regulatory data categories per ingredient. Done by hand it is careful, repetitive reading: open the file, find the statement, retype the value, compare it against the previous revision and against every other document in the folder.
The hardest cost to see is what happens when documents disagree. A conflict stops being one analyst's task and becomes a small cross-functional investigation — R&D, regulatory, quality, and operations reading the same paragraph against each other, at times drawing in senior management. RCMA's baseline was set with that burden on the table.
Stated assumption
The savings baseline — 70 minutes of analyst work and $70 of loaded cost per document — was set by RCMA consensus, prompted by a discussion that surfaced the hidden cost of those cross-functional investigations, which often drew in senior management. It is published as an assumption, not a measurement; every time and cost figure on this page rests on it.
Criteria before evidence.
The acceptance criteria were written into the Pilot Plan in March 2026, before a single pilot document was processed.
Targets and lower mandatory floors were anchored on 18 development-sample documents, so the thresholds could not be chosen to flatter the result. That left the pilot one question to answer: does performance hold at production volume, on documents the system has never seen?
The four weeks ran as four sprints. Documents were processed Monday to Thursday; analysis and refinement happened Friday to Sunday. Two RCMA users met a stable system every working morning — improvement happened on weekends, not in their laps.
The pilot ran with the rigor of a Six Sigma validation process. Acceptance criteria were defined before development began, every result was traceable to its source, and everything was backed by evidence.
The Research Council of Makeup Artists Inc.
SIPOC · Manual ingredient intelligence workflow · improved by Elara
- Ingredient suppliers / technical contacts
- RCMA R&D team
- RCMA regulatory / business stakeholders
- Supplier ingredient documentation
- Existing ingredient records
- RCMA data-category schema
- RCMA business rules / interpretation context
- Excel-ready ingredient intelligence record
- Source references & supporting statements
- Gap / conflict flags
- Supplier follow-up needs
- Updated ingredient database
- R&D
- Regulatory / compliance
- Product development
- Operations / supplier-facing users
Elara flags. People decide.
Elara is an advisory system. It reads, extracts, reconciles, and assesses completeness; it does not make regulatory decisions. Human review precedes any formal use of its output, and every record it produces carries that statement on its face.
Ingest
Supplier documents arrive in batches and are classified by type before anything is read for content.
Extract
Each regulatory data category is captured with its source document and the exact supporting statement attached.
Reconcile
Compatible data is merged; conflicts across documents and against the prior record are surfaced with the reasoning shown.
Assess
Completeness is scored per category, and gaps, conflicts, and compliance-critical items are flagged for follow-up.
What one field looks like
Illustrative · synthetic data
Seven criteria, set in advance. Seven results.
Critical-to-quality criteria · Elara pilot, RCMA Makeup
All targets met
| Criterion | Mandatory | Target | Result | Status |
|---|---|---|---|---|
| Holistic extraction accuracy Every data category assessed on every document — including the cases where the correct answer is that the source says nothing. | ≥ 90% | ≥ 95% | 99.3% n 3,480 | Met |
| Capture extraction accuracy Values that are stated in the document, captured correctly. | ≥ 85% | ≥ 90% | 96.0% n 651 | Met |
| Document classification Each document assigned the right type before its content is read. | ≥ 90% | ≥ 95% | 95.0% n 120 | Met |
| Reconciliation accuracy Matches and conflicts across documents and prior records resolved correctly, with the reasoning shown. | ≥ 85% | ≥ 90% | 93.8% n 209 | Met |
| Task evaluation accuracy The completeness judgement the system makes per task, checked against the reviewers. | ≥ 85% | ≥ 90% | 98.6% n 355 | Met |
| Cycle time per document Wall-clock processing time from ingestion to updated record. | — | < 5 min avg | 145.9 s n 140 | Met |
| Data-integrity failures Any loss, corruption, or mis-assignment of a value during processing. | None permitted | — | None observed in 140 docs | Met |
a Holistic accuracy counts the cases where the correct answer is that the source says nothing — the check that the system does not invent values.
b Figures as recorded in the Elara Pilot Final Report, issued May 2026, covering the window 23 March – 16 April 2026.
In Sprint 2, classification fell below its floor.
Sprint 2 classification
Confirmation run · n 20 × 2
Document classification accuracy dipped to 89.1% in the second sprint — under its mandatory floor.
The root cause was ambiguous handling of signator authority on supplier regulatory statements: the system was not reliably distinguishing whose voice a document represented, and mis-classified those documents as a result.
The fix was a designed series of experiments, not a patch. Candidate prompts, models, and parameter settings were screened systematically against the set of samples that had failed, and the best-performing combination was confirmed against that same set before returning to production volume.
We read the recovery in later sprints with a healthy reservation: those figures alone cannot prove the fix, because later document batches may simply have been easier. The screening against the failing set is the evidence; the sprint trajectory is only consistent with it.
Long term, the outcome is stabilized the way a process should be: a control phase in production — human-in-the-loop review and targeted spot-checks — holds the gain rather than a one-time fix.
Audited, not self-reported.
The accuracy figures are not model self-assessments. Every data point was scored by client reviewers and audited line by line, and nothing was counted until reviewer and auditor agreed. AI-assisted evaluations were themselves manually audited, and the full trail was retained.
The audit pass was performed by Smart-Suited Tech alongside RCMA's reviewers, so we do not describe it as independent. We describe it as documented, and the record is available on request.
The results spoke for themselves: 99.3% accuracy across 140 of our supplier documents.
The Research Council of Makeup Artists Inc.
I would highly recommend them to any organization looking for an experienced, professional AI development team.
The Research Council of Makeup Artists Inc.
Approved for production.
All mandatory floors cleared. All targets met. No data-integrity failures observed, and no critical defects open at close.
RCMA approved production deployment. Elara now runs in RCMA's own cloud tenant: RCMA controls the infrastructure and the spend, with US data residency configurable. Monitoring continues through the human review step and targeted spot-checks — the same review the pilot was built around.
Delivered by Smart-Suited Tech, then operating as Automa Services LLC.
