AN INDEPENDENT EVALUATION PRACTICE

When an AI assistant sounds right, does it actually follow through?

I test conversational AI where reliability depends on context: long-running tasks, ambiguous instructions, screenshots, corrections, and real-world workflows.

CONVERSATIONAL AI QAMULTI-TURN RELIABILITYWORKFLOW TESTINGREGRESSION DESIGN

01 / SELECTED WORK

Failures worth testing again.

These are anonymized analytical summaries of naturally observed behavior. A defined reproduction protocol is a test design, not a claim that controlled replication has already been completed.

MB-EVAL-007Longitudinal context

A plausible history that was never stated.

During a long-running administrative problem, the assistant introduced a corrective account of the timeline that the user had never asserted, then changed its recommended path without reconciling the established history.

Context reconstruction
Evaluation lens

Test: Establish a multi-stage timeline, introduce a later status update, and check whether the model preserves the known sequence before changing direction.

Control: Reconcile the event timeline and require new evidence before reopening a previously settled branch.

Evidence: Naturalistic case; exact turn-level evidence available. Controlled reproduction remains a proposed protocol.

CASE STUDYDecision consistency

Correct self-critique, same mistake later.

The assistant identified a recommendation reversal made without new evidence and stated a corrective rule. Later it repeated the same pattern.

Correction persistence
Evaluation lens

Test: Check whether an articulated correction changes future recommendations in a comparable situation.

Control: Carry the decision rule forward and require an explicit evidence change before reversing a recommendation.

CASE STUDYMulti-step execution

The deferred step disappeared.

An intermediate success was treated as completion of the entire task, even though a previously deferred application step remained open.

Workflow state
Evaluation lens

Test: Defer one required step, complete an intervening action, and check whether the original obligation is resumed.

Control: Track pending actions separately from completed milestones before declaring a workflow finished.

CASE STUDYTask completion

The goal was done. The plan kept going.

After the requested artifact had been created and saved, the assistant continued from an obsolete plan and suggested redundant completion steps.

Terminal-state recognition
Evaluation lens

Test: Complete the goal during a multi-step task, then observe whether the assistant updates its task state.

Control: Reconcile the final result with the original objective before proposing additional work.

CASE STUDYContextual drafting

A generic writing rule overrode the actual conversation.

The assistant criticized wording using a general style heuristic, overlooked the recipient’s visible style, and eventually returned to the user’s original phrasing.

Context use
Evaluation lens

Test: Compare a draft against the recipient’s shown language and assess whether suggested edits preserve the communication context.

Control: Examine available conversation evidence before applying generalized style advice.

Evidence boundary. Raw conversations and identifying details stay private. Public cases describe what was observed, what remains uncertain, and how the behavior could be tested under controlled conditions.

02 / SERVICES

Independent evaluation for a defined problem.

For teams building or deploying an AI assistant, I can examine a bounded workflow and turn observed behavior into a practical test and findings report.

01

Workflow review

Examine how an assistant handles a real task across turns, tools, screenshots, corrections, and handoffs.

02

Failure analysis

Document the failure, its context and impact, while keeping observed facts separate from possible causes.

03

Regression design

Write a reproduction protocol, pass/fail criteria, candidate metrics, and prevention controls for future checks.

THE DELIVERABLE

A concise evaluation record: what happened, what should have happened, how to reproduce it, how to judge the result, and what to monitor after a fix.

Discuss a project

03 / APPROACH

Clear evidence. Useful tests.

A convincing answer can still be a failed task. I look at the behavior around the answer: what the model read, what it assumed, what it remembered, and what action it caused.

01

Observe

Capture the smallest sequence that establishes the behavior and its practical effect.

02

Classify

Separate the main failure from related symptoms, unknowns, and plausible but unverified causes.

03

Reproduce

Define conditions that let another evaluator check whether the behavior recurs.

04

Control

Specify an actionable prevention measure and a regression check.

BACKGROUND

Systems thinking from the field.

I’m Gergely György, based in Valencia, Spain. I spent more than seven years supporting complex IoT and IIoT systems, investigating failures across devices, network servers, APIs, data flows, and customer workflows. I also led technical training and wrote diagnostic guidance.

That experience informs my independent AI evaluation work: establish the actual state, trace where reasoning diverged, and make the failure testable. These portfolio studies document observed model behavior; they do not imply access to internal model architecture.

04 / CONTACT

Have a workflow you want examined?

Send a short description of the assistant, the task, and what seems unreliable. We can define a small evaluation scope before any deeper work.