Arnaldo Sepulveda

Arnaldo Sepulveda

AI Engineer & Builder working across production AI, conversational AI, RAG, evaluation, and enterprise contact-center systems.

I bring 12+ years of enterprise systems engineering across contact-center, cloud, integration, migration, and production incident work at Genesys. Since late 2024, I have applied that foundation hands-on to AI engineering with Python, FastAPI, PostgreSQL, pgvector, hybrid retrieval, evaluation, and local inference. My work focuses on connecting AI capabilities to the operational systems, workflows, evidence, and debugging practices required to make them useful in real environments.

01  Current work

Hands-on AI engineering across retrieval, conversational workflows, and evaluation

Keystone Applied Intelligence is where I build and test reference implementations for conversational AI, RAG, authorization-aware retrieval, agent workflows, and evaluation. The work spans Python and FastAPI APIs, PostgreSQL full-text search, pgvector, hybrid retrieval, deterministic reranking, evidence thresholds, local model serving, OpenTelemetry tracing, and debugging.

These public projects are separately composed engineering instruments, not one demonstrated production runtime. They let me test concrete mechanisms and retain internal evaluation evidence without treating one implementation or passing run as universal validation. Runtime governance is a secondary specialization within that broader AI systems work.

Current engineering mechanisms
  • Python APIs with FastAPI
  • PostgreSQL FTS + pgvector retrieval
  • RAG with evidence thresholds
  • Authorization-aware retrieval
  • Conversational and agent workflows
  • Evaluation and OpenTelemetry tracing
EngageConversational
Conversational-agent reference implementation with a served single-orchestrator path, evidence gating, escalation mechanisms, and optional coordination work outside the default route. keystone-engage
CounselRetrieval
Authorization-aware retrieval reference implementation with query-time role, classification, and client-isolation predicates plus evidence-backed response controls. keystone-counsel
VerifyEvaluation
Standalone HTTP evaluation harness for compatible endpoints, using profiles, deterministic assertions, local results, and run metadata. keystone-verify
LedgerEvidence
Retained internal evaluation artifacts and lineage, including baselines, remediations, and selected failing-to-passing history. keystone-ledger

Secondary research: Governed Execution. Its bounded Track A reference implementation, Runtime Validity, studies controlled process-local authority change, revalidation behavior, and transition evidence. It does not represent external revocation or production authorization. runtime-validity on GitHub.

02  What I’ve learned

Positions earned by implementation

The knowledge may already exist; the hard part is making it usable

In enterprise support, the answer often already exists somewhere: a prior case, documentation, an engineering discussion, or an experienced person’s memory. The harder problem is retrieving the right knowledge for the right person and context, with citations and authorization boundaries, without forcing another engineer to reconstruct the same reasoning.

Evaluation should be able to embarrass the system, not flatter it

An evaluation that cannot surface failures provides weak evidence. The useful ones are built to surface failure (adversarial access probes, out-of-scope queries, fail-closed cases), and the failing runs get published next to the passing ones. That is where the real design feedback comes from.

Production AI is a systems problem

A useful model is only one component. Production behavior also depends on retrieval, APIs, state, integrations, latency, escalation, observability, failure handling, and the surrounding operational workflow.

Some controls belong outside the prompt

When a requirement must hold regardless of model wording, it often belongs in deterministic system logic: retrieval filters, database predicates, state transitions, refusal rules, or authorization checks.

Local and customer-controlled infrastructure exposes engineering tradeoffs

Local inference makes tradeoffs in latency, model capacity, failure modes, operational ownership, and data boundaries explicit. Those constraints can clarify the architecture without making cloud APIs inherently inferior.

03  Writing

Notes from building AI systems

I write about retrieval, evaluation, enterprise AI engineering, and runtime controls: lessons that emerge from implementation, measurement, and operational experience.

Career note · Applied AI
The knowledge was already there. The system to use it was not.

How repeated enterprise support investigations pushed me toward retrieval, AI-assisted workflows, and systems that make existing organizational knowledge useful when people need it.

Engineering note · Evaluation
What evaluation artifacts taught me that demos never could

Retained failures, bounded claims, and the engineering feedback that a polished demonstration cannot provide.

Career note · Enterprise AI
From Genesys to Keystone: enterprise rigor on the LLM substrate

How contact-center operations, escalation, integration, and incident response shape my approach to AI systems.

Writing index & notebook

04  Background

Enterprise engineering is the foundation, not a previous chapter

I spent 12+ years at Genesys working within enterprise contact-center systems. In Business Applications, I specialized in Knowledge and Knowledge Center, AI and classification systems, Digital Services, Agent Workspace, customer and interaction data, routing, conversational systems, and enterprise integrations. This work covered implementations, go-lives, migrations, high-severity incidents, distributed troubleshooting, customer-facing technical investigation, and production recovery.

I also troubleshot WFM-integrated agent and supervisor workflows and operational statistics, tracing missing or incorrect data across application and data boundaries. I worked directly with customers, product managers, developers, and technical directors on product behavior, supportability, customer requirements, and deployment architecture, including clustered and high-volume customer and interaction-data deployments. I later led the Genesys Cloud CX UI Support Team. My AI engineering work since late 2024 builds directly on that experience.

I hold an MScE in Electrical Engineering from the University of New Brunswick, with a thesis applying machine learning to smart-grid load control.

05  Evaluation evidence & selected artifacts

What is public now

I try to make claims that map to something you can inspect. The internal eval baselines below are published with their methodology and lineage in the ledger. They are commit-bound project evidence, not independent validation.

Governed retrieval, keystone-core/retrieval-v1 (2026-04-11): P@1 0.75, MRR 0.79, adversarial ACL 8/8 blocked, fail-closed 5/6, Alberta OHS safety corpus.
Governed agent evaluation run, keystone-core/agent-v1 (2026-05-20): 186 cases across 12 categories and 558 executions, with 153 strict passes, 33 characterization cases, and 0 strict failures at the evaluated keystone-gov commit. The failing precursor run is published alongside it.

06  Contact

Work I want to do next

I am interested in production AI work where I can combine enterprise systems experience with hands-on AI building: conversational AI, RAG and retrieval, AI-assisted workflows, evaluation, analytics, APIs, and enterprise integrations. I am especially interested in systems used in real customer and operational workflows where outcomes can be measured, investigated, and improved. If that is the kind of problem you are staffing, I would like to hear about it.