Uncategorized

Het Brein van Art Revisionist – AI die Documentatie Analyseert

Van Database naar Denkend Systeem. We bouwden een AI-systeem dat niet alleen informatie opslaat, maar het ook begrijpt, verbanden legt, en tegenstrijdigheden detecteert. Domein-specifieke AI getraind op primaire bronnen.

Post 2: Het Brein van Art Revisionist – AI die Documentatie Analyseert

Nederlands:

Van Database naar Denkend Systeem

In mijn vorige post vertelde ik hoe we alle Valsuani-informatie verzamelden in een gestructureerd systeem. Maar een verzameling documenten, hoe goed georganiseerd ook, is nog geen begrip. We hadden data, maar we wilden inzicht.

De vraag was: kunnen we een AI-systeem bouwen dat niet alleen informatie opslaat, maar het ook begrijpt, verbanden legt, en tegenstrijdigheden detecteert?

Het antwoord: ja. Maar niet op de manier zoals je AI meestal gebruikt.

Het Probleem met Standaard AI

Als je ChatGPT vraagt “Wie stichtte de Valsuani gieterij?”, krijg je het antwoord dat gebaseerd is op zijn trainingsdata – die vol staat met de foute informatie die we juist proberen te corrigeren. AI-systemen zijn briljant in het samenvatten van wat er al bekend is, maar minder goed in het ontdekken dat wat “bekend” is eigenlijk fout is.

We hadden iets anders nodig: een AI dat specifiek getraind was op onze primaire bronnen, niet op de secundaire bronnen die de fout perpetueren.

Kennisbank-Architectuur

We bouwden wat je zou kunnen noemen een “domein-specifieke AI” – een systeem dat:

1. Documenten Analyseert
– OCR-scan van historische documenten (handgeschreven geboorteaktes, getypte zakelijke brieven)
– Automatische extractie van namen, datums, plaatsen, relaties
– Classificatie van betrouwbaarheidsniveau per bron (primair vs secundair vs tertiair)
2. Verbanden Legt
– “Carlo Valsuani overleed in 1886” + “Claude was 10 jaar oud bij vaders dood” = Claude geboren in 1876
– Cross-referencing tussen familiedocumenten en zakelijke archieven
– Tijdlijnconstructie: wie kon waar zijn geweest, wanneer
3. Tegenstrijdigheden Detecteert
– Veilingcatalogus zegt “Marcello stichtte gieterij in 1900”
– Maar geen enkel Frans zakenregister vermeldt “Marcello” voor 1900, 1910, 1920…
– AI markeert dit als “claim zonder primaire bron”
4. Bewijskracht Weegt
– Primaire bron (geboorteakte) = hoge bewijswaarde
– Secundaire bron (museumcatalogus) = middelmatige bewijswaarde
– Tertiaire bron (blog zonder bronvermelding) = lage bewijswaarde

Wat AI Wel en Niet Kan

Hier werd het interessant. AI bleek uitzonderlijk goed in taken die voor mensen saai en foutgevoelig zijn:

AI excelleert in:
– Patronen herkennen over 500+ documenten (“deze naam verschijnt nergens”)
– Tijdlijn-inconsistenties spotten (“persoon X kan niet op twee plekken tegelijk zijn”)
– Citeerketens traceren (“deze 50 bronnen citeren allemaal één verkeerde catalogus uit 1971”)
AI worstelt met:
– Context begrijpen van historische dubbelzinnigheden
– Handgeschreven documenten in slecht Italiaans uit 1880
– Culturele nuances (waarom zou iemand “Marcel” verwarren met “Marcello”?)

De oplossing: menselijke expertise voor context, AI voor schaal.

Het Doorbraakmoment

Ergens in de tweede maand van development gebeurde iets bijzonders. We voerden een nieuw document in – een zakelijke brief uit 1908 waarin stond “Claude Valsuani, fils de feu Carlo” (zoon van wijlen Carlo).

De AI herkende niet alleen de familierelatie, maar merkte ook op:
– “Carlo” hier vermeld als overleden (“feu” = wijlen)
– Sterfcertificaat Carlo: 1886
– Claude zou 10 jaar oud zijn geweest bij vaders dood
– Geen enkel document suggereert dat Carlo “Marcello” als tweede naam had

De AI concludeerde wat mensen over het hoofd hadden gezien: als Carlo de vader was, en Carlo in 1886 overleed, en er geen documenten zijn die “Marcello” noemen, dan is “Marcello” waarschijnlijk een latere verzinning.

Dit was geen magische AI-revelatie. Dit was systematisch redeneren op basis van chronologie en afwezigheid van bewijs. Maar het uitvoeren van deze redenering over honderden documenten? Dat is waar AI schittert.

Technische Stack

Voor de geïnteresseerden, onze setup:

Document Processing: Python + Tesseract OCR voor handgeschreven documenten
Knowledge Base: Gestructureerde database met relaties tussen entiteiten
AI Analysis: Custom GPT-4 fine-tuning op kunsthistorische attributie-logica
Conflict Detection: Regelgebaseerd systeem dat claims tegen primaire bronnen checkt
Timeline Engine: Chronologische consistentie-checker

Van Analyse naar Synthese

De kennisbank werd niet alleen een opslagplaats, maar een actief redeneersysteem. Stel een vraag, en het systeem:

1. Zoekt relevante documenten
2. Weegt bewijskracht
3. Detecteert tegenstrijdigheden
4. Presenteert conclusie met betrouwbaarheidsscore

Vraag: “Wie stichtte de Valsuani gieterij?”
Antwoord: “Claude Valsuani (zoon van Carlo, 1876-1923). Betrouwbaarheid: 98%. Basis: 23 primaire documenten. Conflicterende claims: ‘Marcello Valsuani’ verschijnt in 47 secundaire bronnen zonder primaire onderbouwing.”

Volgende Stap: Meerdere AI’s Loslaten

We hadden nu een AI die onze verzamelde kennis kon analyseren. Maar wat als we niet alleen onze eigen documenten gebruikten, maar het hele internet doorzochten op conflicterende informatie?

Dat werd de volgende fase: meerdere AI-systemen die parallel werkten, elk met een andere invalshoek, elk met verschillende online bronnen. En wat ze vonden was… verrassend.

Daarover meer in de volgende post.

Bekijk het volledige resultaat op Art Revisionist, waar alle analyse transparant gedocumenteerd staat.
Geïnteresseerd in hoe we AI gebruiken voor kennisopbouw? Zie ook ons werk bij Prospergenics.

English:

From Database to Thinking System

In my previous post I told you how we gathered all Valsuani information into a structured system. But a collection of documents, however well organized, isn’t understanding yet. We had data, but we wanted insight.

The question was: can we build an AI system that doesn’t just store information, but also understands it, makes connections, and detects contradictions?

The answer: yes. But not the way you usually use AI.

The Problem with Standard AI

If you ask ChatGPT “Who founded the Valsuani foundry?”, you get an answer based on its training data – which is full of the incorrect information we’re trying to correct. AI systems are brilliant at summarizing what’s already known, but less good at discovering that what’s “known” is actually wrong.

We needed something different: an AI specifically trained on our primary sources, not on the secondary sources that perpetuate the error.

Knowledge Base Architecture

We built what you might call a “domain-specific AI” – a system that:

1. Analyzes Documents
– OCR scan of historical documents (handwritten birth certificates, typed business letters)
– Automatic extraction of names, dates, places, relationships
– Classification of reliability level per source (primary vs secondary vs tertiary)
2. Makes Connections
– “Carlo Valsuani died in 1886” + “Claude was 10 years old at father’s death” = Claude born in 1876
– Cross-referencing between family documents and business archives
– Timeline construction: who could have been where, when
3. Detects Contradictions
– Auction catalog says “Marcello founded foundry in 1900”
– But no French business register mentions “Marcello” for 1900, 1910, 1920…
– AI flags this as “claim without primary source”
4. Weighs Evidence
– Primary source (birth certificate) = high evidential value
– Secondary source (museum catalog) = medium evidential value
– Tertiary source (blog without citations) = low evidential value

What AI Can and Cannot Do

This is where it got interesting. AI proved exceptionally good at tasks that are boring and error-prone for humans:

AI excels at:
– Recognizing patterns across 500+ documents (“this name appears nowhere”)
– Spotting timeline inconsistencies (“person X cannot be in two places at once”)
– Tracing citation chains (“these 50 sources all cite one incorrect 1971 catalog”)
AI struggles with:
– Understanding context of historical ambiguities
– Handwritten documents in poor Italian from 1880
– Cultural nuances (why would someone confuse “Marcel” with “Marcello”?)

The solution: human expertise for context, AI for scale.

The Breakthrough Moment

Somewhere in the second month of development, something special happened. We entered a new document – a business letter from 1908 stating “Claude Valsuani, fils de feu Carlo” (son of the late Carlo).

The AI not only recognized the family relationship, but also noted:
– “Carlo” mentioned here as deceased (“feu” = late)
– Carlo’s death certificate: 1886
– Claude would have been 10 years old at father’s death
– No documents suggest Carlo had “Marcello” as a second name

The AI concluded what people had overlooked: if Carlo was the father, and Carlo died in 1886, and there are no documents mentioning “Marcello”, then “Marcello” is probably a later invention.

This wasn’t a magical AI revelation. This was systematic reasoning based on chronology and absence of evidence. But performing this reasoning across hundreds of documents? That’s where AI shines.

Technical Stack

For those interested, our setup:

Document Processing: Python + Tesseract OCR for handwritten documents
Knowledge Base: Structured database with relationships between entities
AI Analysis: Custom GPT-4 fine-tuning on art historical attribution logic
Conflict Detection: Rule-based system checking claims against primary sources
Timeline Engine: Chronological consistency checker

From Analysis to Synthesis

The knowledge base became not just a repository, but an active reasoning system. Ask a question, and the system:

1. Searches relevant documents
2. Weighs evidential strength
3. Detects contradictions
4. Presents conclusion with confidence score

Question: “Who founded the Valsuani foundry?”
Answer: “Claude Valsuani (son of Carlo, 1876-1923). Confidence: 98%. Based on: 23 primary documents. Conflicting claims: ‘Marcello Valsuani’ appears in 47 secondary sources without primary support.”

Next Step: Unleashing Multiple AIs

We now had an AI that could analyze our collected knowledge. But what if we didn’t just use our own documents, but searched the entire internet for conflicting information?

That became the next phase: multiple AI systems working in parallel, each with a different angle, each with different online sources. And what they found was… surprising.

More on that in the next post.

View the complete result at Art Revisionist, where all analysis is transparently documented.
Interested in how we use AI for knowledge building? See also our work at Prospergenics.

Publicatiedatum: [DAY 2]
Tags: Artificial Intelligence, Knowledge Systems, Art Authentication, Machine Learning, Valsuani
Categorie: Projects

Terug naar overzicht
ENNL