Skip to content

Build in Public

This technical log follows, as we go, the components we are developing. It shows the real state of the work (what works, what remains to be done, what the measurements say) rather than a showcase. Demonstrations rely on public or fictitious data, labelled as such.

All demonstrations are carried out in test environments, without any client data.

The format of each entry

Scoping

Perimeter, technical choices and test set are documented before any development starts.

Prototype

A first working version, first measurements, difficulties identified and addressed.

Demonstration

Demonstration on the test set, measured results and limitations stated plainly.

  • Mock-up on a synthetic corpus

    Document search mock-up on ten contract clauses

    A document search mock-up built on ten synthetic contract clauses (confidentiality, personal data, subcontracting, reversibility…), a few lines each. The search compares the term frequencies of the question with those of each clause; it uses no chunking, no vector representation model and no answer generation.

    Scoping
    Ten synthetic clauses; fifteen questions, each linked to the clause that answers it and to the terms the answer must contain.
    Prototype
    Search by term-frequency similarity, on the whole clause; an automatic evaluation checks the clause retrieved and the presence of the expected terms.
    Demonstration
    The right clause comes first for 13 questions out of 15. In-house composite score (0.6 × retrieval + 0.4 × presence of the expected terms): 0.867. This is not a RAGAS evaluation.

    Lessons

    • Both failures concern questions whose answer is in the corpus (duration of the confidentiality obligation, information to be documented under the GDPR): the term-based search did not bring up the right clause.
    • The question set contained no out-of-corpus question: the ability to refuse was not measured.
    • A useful evaluation checks both the document retrieved and the content of the answer.

    SourcePublic code, data and report: oppdrag-public-pocs, folder poc-001-rag-juridique.

  • Fictitious architecture, deliberately misconfigured

    Configuration audit of a fictitious cloud architecture

    An audit script applied to the description of a fictitious AWS architecture, deliberately misconfigured for the exercise: accounts, storage and network rules. No connection to an AWS account: everything is local and fictitious.

    Scoping
    Scope: identity management, storage, network rules; seven control rules, each with a severity and a remediation.
    Prototype
    Full audit: 23 findings, 5 critical, 6 high, 6 medium and 6 low, ranked by severity.
    Demonstration
    Final report, with a proposed remediation for each finding. The fixes were not applied.

    Lessons

    • The five critical findings concern identity management (administrator accounts without multi-factor authentication) and network exposure (SSH access and a database open to the internet).
    • An audit should also verify the fixes, not only detect: that step is missing from this exercise.

    SourcePublic code, data and report: oppdrag-public-pocs, folder poc-004-audit-cloud-aws-fictif.

  • Trials on the delivered product, 5 and 6 October 2026

    What the demonstration revealed in Knowledge AI

    For this site's demonstration, Knowledge AI was queried, from its main branch, on a public extract of sixteen articles of the French Labour Code. The trials confirmed what they had to confirm, and revealed defects that we recorded rather than worked around.

    Scoping
    Sixteen articles of the French Labour Code (Légifrance, Open Licence), in three PDFs; three questions a company director would ask, one of them with no answer in the corpus.
    Prototype
    Indexing on a dedicated database; measurement of the evidence check and reconstruction of the context passed to the model, without generation.
    Demonstration
    The unanswerable question was refused five times out of five. The other two were wrongly refused; rephrased in the wording of the text, they pass the check, but their generation, on CPU only, takes more than three minutes.

    Lessons

    • A protected PDF failed without explanation: it is now quarantined with an understandable reason. A damaged PDF must be handled likewise (issue 196).
    • The evidence check compares words, not meaning: “contractor” instead of “outside company” is enough to have a legitimate question refused, with 10% of its terms found for a 25% threshold (issue 197).
    • The refusal does not say why the tool is silent, and suggests a reference foreign to the corpus (issue 198).
    • Chunking can separate an article's heading from its text: 6 articles out of 16 are split, and a passage can cite an article without containing its text (issue 199).
    • The audit trace names the model declared in the configuration, not the one that answered: this is the highest-priority defect, because it affects the audit trail (issue 200).

    SourcesRedesign journal (sprint 5) and measurement scripts; issues 194 to 200 of the Knowledge AI repository, private, presented during a demonstration.

Logbook

Longer analyses (architecture choices, methods, lessons learned) are published in the logbook.

Open the logbook

A comparable need?

These demonstrations illustrate how we work. Tell us about your project: we will examine it with the same rigour.