Technical Partner
I help founders decide what is technically viable before they commit time and budget to an AI product.
Abstract
I work with founders who face a costly technical decision before they have a full engineering team. We define the question, inspect whatever already exists, then run a small test that can change the answer. The papers on this site show how I work. They document systems I built and failures I found, including conclusions I had to narrow. I can assess feasibility and implementation risk. I cannot validate demand or product-market fit.
Keywords: systems evaluation; conversational memory; language models; technical feasibility
Scope of the work
Model behavior is only part of the problem. Memory, permissions, bad inputs, latency, cost, and recovery often decide whether a product can be operated. I review those constraints before a team commits to an architecture.
I do not offer customer discovery, market sizing, pricing, sales, or product-market fit work.
Independent research
Three full papers are available here. The first defines a tenant-isolation protocol for conversational memory. The second audits a preserved 46-case run of a restaurant agent. The third documents an extraction pipeline, reproduces its 569 software tests, and states the benchmark that still has to be run. These are independent working papers. None has been peer reviewed.
WP-01 · Working paper · 2026
Tenant-Scoped Long-Term Memory in Conversational Agents: Architecture, Threat Model, and an Executable Evaluation Protocol
Can a multi-tenant conversational agent recover durable user facts without exposing facts from another tenant, and can that claim be tested with an inspectable protocol?
Evidence status: Protocol and implementation inspected; no public live-run artifact
TR-01 · Technical report · 2026
Failure-Focused LLM-as-a-Judge Evaluation for a Transactional Conversational Agent: A 46-Case Repository Study
Can a category-based model judge expose actionable failure clusters in a multi-turn transactional agent, and what must be added before its scores are reliable evidence?
Evidence status: Saved 46-case run inspected; provenance gaps and test-design limits documented
RP-01 · Research protocol · 2026
Provenance-Constrained Multi-Stage LLM Extraction from Heterogeneous Product Catalogs: System Design and Benchmark Protocol
Does separating document understanding, source-row reconstruction, and product normalization improve reliability over a single-pass extraction prompt?
Evidence status: Architecture and 569 software tests reproduced; extraction benchmark still unexecuted
Ways to work together
Bring one decision that could change the architecture, cost, or feasibility of the product. I review the material, agree with you on what would count as an answer, and return a recommendation you can use. The USD 120 session fee is credited toward a sprint if I accept it.
| Engagement | Scope and output | Timing | Fee |
|---|---|---|---|
| Technical Decision Session | About 10 minutes of pre-work, one 75-minute session, and a one-page Technical Decision Brief delivered within 24 hours. | 75 minutes | USD 120 for the first five clients; then USD 180. |
| Technical Hypothesis Sprint | Two sessions, one risky technical hypothesis, a fixed-scope proof or technical spike, limited feedback, stated risks, and a 30-day technical plan. | 7 days | USD 550 for the first three cases; then USD 850. |
How I reach a recommendation
First we write the decision in plain language and name the result that would change it. I separate facts from assumptions, then choose a test small enough to run but strong enough to rule an option in or out. The final brief states the recommendation and where it stops being reliable.
QuestionWhat is knownAssumptionTestVIABLE / CONDITIONAL / NOT JUSTIFIED
Boundaries and fit
- One engagement addresses one primary technical question. One sprint can be active at a time.
- The work can assess technical feasibility and implementation risk; it does not validate demand, pricing, distribution, sales, or product-market fit.
- No engagement guarantees funding, revenue, adoption, production success, or any commercial result.
- Complete MVP delivery, open-ended development, and indefinite support are outside scope.
- Logistics, last-mile, routing, fleet, dispatch, delivery management, delivery tracking, and adjacent competitive work are excluded.
Submit a technical question
Tell me what you are trying to decide and what material already exists. I use the form to check fit and conflicts. It does not book or charge anything.
Open the application
All questions are required unless marked optional.
References and contact
- Complete research record
- Conversational-agent evaluation repository
- Document-extraction pipeline repository
Pablo Cardozo · Technical Partner · Independent researcher · Buenos Aires, Argentina