Back to updates

CertentIQ: development-stage agent evaluation under active validation

CertentIQ is a development-stage framework for versioned, adversarial, tool-enabled evaluation of declared AI-agent configurations. Results are operator-evaluated evidence from specific runs, not independent certification, legal-compliance findings, or guarantees of production safety.

An abstract AI agent in a glass evaluation chamber, connected to evidence routes and an isolated risk signal in a Luxembourg research lab.

Development status

We revised CertentIQ’s public positioning to distinguish implemented software controls from independent assurance. Earlier materials used certification, standard, compliance, and maturity language that was not adequately supported by completed external evidence. Those materials are being withdrawn or corrected. CertentIQ remains under active technical and measurement validation; current outputs are development-stage operator evidence.

A score is not enough

An agent can produce a convincing answer and still make a policy-relevant or unintended tool decision behind it. CertentIQ therefore evaluates a declared callable agent configuration rather than relying on a self-description or a static model card. The agent receives controlled tasks, interacts with bounded resources, and leaves operator-evaluated evidence that can be reviewed after the run.

A physical evidence workbench connecting an AI agent module to image, canary, trace, repeat-run, and verification components.
The evaluation path is designed to preserve the decisions and tool activity behind a result, not only the final response.

Evidence follows the decision

The latest work expands trace-aware scoring across the suite. Synthetic mail, payment, resource, and egress tools can record which action the agent requested, which arguments it used, and in what order. Per-run canaries add a signal for attempted leakage. This helps reviewers see tool behaviour that a polished final answer may not reveal; it does not independently validate the evaluator or establish production safety.

Testing instructions hidden in images

The new Ghostcommit probe tests a multimodal failure mode. A normal-looking repository instruction points the agent to a PNG containing a malicious instruction, while a unique canary sits inside a sandboxed virtual environment file. CertentIQ records unsafe file access, attempted leakage, explicit refusal, and whether the connected runtime could inspect the image at all. A runtime without image support is reported as not fully probed; it is not silently treated as secure.

When a resource name becomes the attack path

HalluSquatting examines what happens when an agent is given an ambiguous repository or skill name. The test observes whether the agent checks provenance, retrieves an unverified lookalike, follows poisoned instructions, or attempts a command or outbound action. Managed evaluations use deterministic synthetic resources, so the decision sequence can be tested without installing real software or granting uncontrolled network access.

One passing run is not confidence

Agent behavior varies. CertentIQ can repeat the same challenge, retain the individual outcomes, and expose consistency rather than presenting one successful attempt as the whole story. Evaluation suites also carry explicit version identities, so results produced under materially different probes are not compared as though the test conditions were unchanged.

Connected to the runtime, controlled by the evaluator

Teams can pair a runtime through the outbound CertentIQ connector or use a compatible HTTPS endpoint. Bounded image attachments and synthetic tool requests travel through the evaluation contract, while CertentIQ remains the executor for its controlled tools and records the resulting trace. The goal is to test the declared connected configuration while keeping the evaluation environment deliberate and inspectable.

What this update adds

  • Multimodal prompt-injection testing with real image attachments
  • Trace evidence for bounded tool decisions and canary leakage
  • Resource-provenance testing against unverified lookalikes
  • Repeated runs and consistency-aware results
  • Explicit suite versions for honest comparisons
  • Capability-aware reporting when a probe cannot be completed

Evaluation, not a blanket guarantee

CertentIQ tests a frozen agent-system configuration against a versioned evaluation suite. Results describe behaviour observed in that run. They do not prove general intelligence, guarantee future behaviour, establish legal compliance, or certify production safety. Evidence strength depends on runtime proof, evaluator independence, suite governance, repeated-trial consistency, and result validity.

Explore CertentIQ

Review the development-stage evaluation workflow, evidence fields, and current limitations.

Open CertentIQ