Project / applied AI reliability
Evaluate AI agents as operational software, not as a polished demo.
Experiments with tools, retrieval, fallbacks, observability and adversarial conversations for more reliable agents.

The real problem
A convincing demonstration can still fail when a source is missing, a tool times out or a user asks something outside the intended boundary. Reliability begins by making those states explicit.
What the lab exercises
Scenarios cover source grounding, tool permissions, structured outputs, handoff, retries and logs that allow a failed conversation to be understood. Evaluation is attached to representative tasks rather than subjective impressions.
Responsible claims
Local evaluation cannot prove every production conversation or replace monitoring. Model behavior, provider availability, cost and source quality can change, so operation requires ongoing measurement and safe fallback.
WhatsApp / (85) 98596-3329