Research on Production AI Systems
Abstract
This record separates research claims from engineering artifacts. Working papers define questions and evaluation protocols; technical reports document implemented systems and bounded evidence; research protocols state what still needs to be measured. None of the current works is presented as peer reviewed.
Keywords: AI evaluation; conversational memory; agent systems; structured extraction; production AI
1. Current work
Each entry states its evidence status and limitations. A polished interface is not treated as a result, and repository-reported measurements are labelled as such.
WP-01 · Working paper · 2026
Evaluating Long-Term Memory Isolation in Multi-Tenant Conversational Agents
This working paper defines a reproducible evaluation protocol for conversational memory under tenant isolation. The system combines short-term thread state, vector memory, and graph-grounded documentary evidence. The available artifact creates two synthetic tenants, records distinct facts for each, probes recall, checks cross-tenant leakage, and measures latency. No quantitative result is claimed until a versioned run artifact is released.
- Status
- Evaluation protocol available; public result artifact pending
TR-01 · Technical report · 2026
Judge-Driven Evaluation of a Transactional Conversational Agent
This technical report examines an evaluation harness for a restaurant assistant that must answer questions, manage multi-turn orders, handle payments, resist unsafe inputs, and escalate to a human. A separate judge model scores accuracy, actionability, completeness, relevance, and tone on a 0 to 100 scale, while the harness records latency, tokens, traces, and category-level statistics. A dated repository report records an increase from 19 of 46 passing cases to 31 of 46 after a partial implementation. These figures are repository-reported and have not been independently rerun for this publication.
- Status
- Public implementation and repository-reported results
RP-01 · Research protocol · 2026
A Benchmark Protocol for Multi-Stage LLM Extraction from Heterogeneous Product Catalogs
This research protocol turns an implemented three-stage catalog ingestion system into a falsifiable comparison. The system first extracts document structure, then reconstructs source rows, and finally normalizes products in bounded batches with schema validation and provenance checks. The proposed benchmark compares this decomposition with a single-pass baseline on field accuracy, row coverage, provenance integrity, review rate, latency, and cost. No comparative result is currently claimed.
- Status
- Implemented system; controlled benchmark not yet run
2. Publication standard
A work moves from protocol to report only after its question, dataset, baseline, environment, metrics, raw outputs, and limitations can be inspected. A report becomes a submission candidate only after reproduction and external review.
3. Disclosure
AI systems were used during software development and editorial preparation. Pablo Cardozo remains responsible for the claims, source selection, interpretation, and corrections. Authorship will be assigned according to substantive contribution, not repository ownership alone.
Pablo Cardozo · Independent Applied AI Researcher · Buenos Aires, Argentina