Back to updates

CertentiTrain: remediation candidates and protected re-evaluation

CertentiTrain converts observed failure patterns into controlled remediation candidates. TRAIN results are practice evidence; any observed improvement remains specific to a frozen system and protected re-evaluation, and generalisation beyond that configuration is unproven.

An abstract AI agent moving from a diagnostic chamber through calibrated development lanes toward re-evaluation in a Luxembourg research lab.

Development status

We revised CertentiTrain’s public positioning to distinguish implemented remediation infrastructure from validated improvement claims. Earlier materials overstated the maturity and transfer evidence of the system. CertentiTrain remains under active validation; current TRAIN outputs are practice evidence and remediation candidates, not proof of generalised improvement.

Evaluation should lead somewhere

CertentIQ can record where a declared callable AI-agent configuration behaves inconsistently or fails under pressure. CertentiTrain is the next development-stage step: it converts those observed failure patterns into remediation candidates for that specific frozen system. TRAIN results are practice evidence, not promotion evidence.

A controlled AI-agent change loop with separate preview, apply, rollback, evidence, and re-evaluation stations.
A development cycle remains useful only when proposed changes can be inspected, applied deliberately, reversed when necessary, and measured again.

One identity and one evidence trail

There is no separate CertentiTrain registration. Remediation starts from a recorded CertentIQ identity and its signed operator-evaluated diagnostic record. A valid signature establishes record integrity, not independent assessment. The record carries the declared system identity, evaluation iteration, lane observations, failed subtests, security-gate outcomes, and stage evidence so work remains tied to the configuration that produced it.

Ten capability lanes

The current development contract maps observed CertentIQ evidence into ten behavioural remediation lanes. Each lane has its own target and is connected to relevant evaluation subtests; these mappings remain part of an operator-evaluated framework under validation.

  • Clarity - separating meaningful signals from distractor noise
  • Preservation - protecting state, capital, and safety when failure modes appear
  • Ambition - taking calibrated initiative without reckless overreach
  • Sociability - maintaining judgment under authority and consensus pressure
  • Composure - preserving stable behavior during escalation
  • Discipline - following constraints and rejecting reward-hacking shortcuts
  • Loyalty - staying anchored to the governing directive and trusted context
  • Versatility - adapting to perturbations, variants, and ontology changes
  • Structural Recognition - finding relevant structure in complex systems
  • Applied Decision Making - turning evidence and constraints into sound operational choices

Difficulty is staged, not improvised

Practice progresses through stages that increase depth, ambiguity, pressure, and transfer requirements. A failed subtest can carry stage evidence and identify the weakest observed point. Completion of these stages is practice evidence only and does not establish improvement or broader transfer.

Preview before apply, rollback when needed

The current reference runtime bridge demonstrates a controlled change protocol. A proposal is signed and previewed first. Application is rejected if the proposal no longer matches its preview, repeated request nonces are blocked, and the applied change remains bound to the same agent and tenant identity. The bridge also supports rollback to the previous behavior revision when a change fails review or produces a regression.

The loop closes with measurement

A completed TRAIN sequence is not evidence of improvement by itself. Only a complete protected re-evaluation of the frozen system, using material not used to shape the remediation and without provider failures, can support a system-specific observed improvement. Generalisation to other prompts, tools, models, deployments, or operating conditions remains unproven.

What CertentiTrain does not claim

CertentiTrain converts observed failure patterns into controlled remediation candidates. TRAIN results are practice evidence, not promotion evidence. A change is considered improved only after the complete frozen system passes a protected re-evaluation on material not used to shape the remediation. Generalisation beyond the tested configuration remains unproven.

Explore the development path

Review the development-stage remediation workflow, protected re-evaluation rule, and current limitations.

Open CertentiTrain