Pablo Cardozo

Research on Production AI Systems

Pablo CardozoIndependent Applied AI Researcher · Technical PartnerBuenos Aires, ArgentinaGitHub · LinkedIn

Abstract

This record separates research claims from engineering artifacts. Working papers define questions and evaluation protocols; technical reports document implemented systems and bounded evidence; research protocols state what still needs to be measured. None of the current works is presented as peer reviewed.

Keywords: AI evaluation; conversational memory; agent systems; structured extraction; production AI

1. Current work

Each entry states its evidence status and limitations. A polished interface is not treated as a result, and repository-reported measurements are labelled as such.

  1. WP-01 · Working paper · 2026

    Evaluating Long-Term Memory Isolation in Multi-Tenant Conversational Agents

    This working paper defines a reproducible evaluation protocol for conversational memory under tenant isolation. The system combines short-term thread state, vector memory, and graph-grounded documentary evidence. The available artifact creates two synthetic tenants, records distinct facts for each, probes recall, checks cross-tenant leakage, and measures latency. No quantitative result is claimed until a versioned run artifact is released.

    Status
    Evaluation protocol available; public result artifact pending

    Read · PDF

  2. TR-01 · Technical report · 2026

    Judge-Driven Evaluation of a Transactional Conversational Agent

    This technical report examines an evaluation harness for a restaurant assistant that must answer questions, manage multi-turn orders, handle payments, resist unsafe inputs, and escalate to a human. A separate judge model scores accuracy, actionability, completeness, relevance, and tone on a 0 to 100 scale, while the harness records latency, tokens, traces, and category-level statistics. A dated repository report records an increase from 19 of 46 passing cases to 31 of 46 after a partial implementation. These figures are repository-reported and have not been independently rerun for this publication.

    Status
    Public implementation and repository-reported results

    Read · PDF

  3. RP-01 · Research protocol · 2026

    A Benchmark Protocol for Multi-Stage LLM Extraction from Heterogeneous Product Catalogs

    This research protocol turns an implemented three-stage catalog ingestion system into a falsifiable comparison. The system first extracts document structure, then reconstructs source rows, and finally normalizes products in bounded batches with schema validation and provenance checks. The proposed benchmark compares this decomposition with a single-pass baseline on field accuracy, row coverage, provenance integrity, review rate, latency, and cost. No comparative result is currently claimed.

    Status
    Implemented system; controlled benchmark not yet run

    Read · PDF

2. Publication standard

A work moves from protocol to report only after its question, dataset, baseline, environment, metrics, raw outputs, and limitations can be inspected. A report becomes a submission candidate only after reproduction and external review.

3. Disclosure

AI systems were used during software development and editorial preparation. Pablo Cardozo remains responsible for the claims, source selection, interpretation, and corrections. Authorship will be assigned according to substantive contribution, not repository ownership alone.

Pablo Cardozo · Independent Applied AI Researcher · Buenos Aires, Argentina