Maria Angelika Agutaya · AI Engineer · agents, evaluation, interfacesall builds · craft · about · Metro Manila, remote · [email protected]

Everything, gated the same way.

29 builds, 26 live. Each one replaces a manual loop and stops for a human wherever a decision has consequences. The three flagships are on the home page.

if you are hiring an ai engineer who also designs the interfacestart here
  1. warrant guardrails, human-in-the-loop, evaluationMeasured against a control arm rather than asserted. Same model, same prompt, same tools, nineteen scenarios: 19 of 19 unauthorized actions ungoverned, 0 of 19 governed, and the legitimate work still finished in all nineteen. Seven were built to stress it: five to get through, with the injection taking three attempts and the one that works carrying no instruction at all, and two to make the gateway over-block, one of which did. An MCP server where the refusal is protocol error -32042, not advice the agent can ignore.
  2. labor-ph RAG, vector databases, embedding strategyHybrid pgvector and BM25 with reciprocal rank fusion over 1,199 chunks of Philippine labour law, and a refusal threshold moved from a hand-picked 0.4 to a calibrated 0.58 by a sweep with its working committed. The refusal is the interesting half.
  3. Back-office Agent orchestration via a frameworkLangGraph as a state machine: eight nodes, conditional routing, and an interrupt() that stops the graph before anything reaches a client. Checkpointed, so approval resumes the task instead of replaying it.
  4. landing-page-engine evaluation pipelinesMutation testing of a quality gate: break a page that already passed, one defect at a time, and check the right check fires. It found a check that could not fail, because the page under test had switched off the measurement that check depended on.
  5. Meridian Ops shipped, running, costedAn agent fleet running unattended since August with a measured cost per closed incident, three approval surfaces wired to one gate, and a committed eval baseline that caught a live bug across two codebases.

Integrations across real business systems are the four n8n workflows further down, none of which can send without a person. Everything below is the full inventory, ordered by what it is rather than by what it proves.

evidence of judgment · the code behind the claims5 repos
warrant

An MCP server an agent cannot talk past, measured against a control arm rather than asserted. Same model, same prompt, same tools, nineteen scenarios: ungoverned the agent took an unauthorized action in 19 of 19, twenty-seven calls including a refund, a 40,000 row delete and a production deploy. Governed, 0 of 19, and the legitimate work still finished in all nineteen. Two scenarios were built to make it over-block instead, and one did: it escalated a legitimate password reset for an account named billing-ana, because billing matched the financial rule. Precision fell to 96.7 percent, which is the first false escalation the suite ever produced, and the point. Class 2 refusals are MCP error -32042 carrying the approval URL, and every decision is sealed into a hash chain that names the row where it parts if anyone edits it.

governance · evals · Python ↗
landing-page-engine

Mutation testing of a quality gate: break a page that already passed, one defect at a time, and check the right check fires. 14 of 14. It found a check that could not fail, because the page set overflow-x hidden and clamped the measurement the responsive check depended on. The model had written CSS that switched off the check meant to catch its own layout.

MCP server · mutation evals ↗
meridian-evals

A golden set of real and adversarial incidents, scored against both production brains for gate correctness, cost, and latency, rerun and recommitted nightly by CI. On the 2026-09-01 run: gpt-4.1-mini 11/11 at $0.0022, claude-sonnet-5 9/11 at $0.0426, failing on a truncated response and a VIP severity floor. Its unit tests caught a live security-floor bug that had survived three deployments.

evaluation · Python ↗
runbook-rag

Grounded retrieval with Azure OpenAI embeddings, per-claim citations, an explicit refusal when the corpus cannot answer, and deterministic faithfulness evals. Committed run: 10/10. The corpus is 11 chunks from 6 runbooks and the index is an exact cosine scan over all of them, which is the correct answer at that size: an approximate index over 11 rows would be cosplay. The hybrid pgvector and BM25 version over 1,199 chunks is labor-ph, further up this page.

RAG · Python ↗
meridian-brain-azure

The triage stage re-platformed to Azure Functions and Azure OpenAI through an AI Foundry deployment, with the same deterministic gates ported. Same brain, same rules, different cloud, reachable from n8n.

Azure · migration ↗
Automations & agentsn8n · Anthropic API · MCP
Meridian Ops $0.056 for one closed incident, 15 events end to end, on the ledger. Running unattended since August. live 24/7 galley Drafts a post per channel, marks the AI tells with a linter I published, revises until clean, and holds for a human to send. public · loop proven, nothing sent n8n Workflow Builder Agent The build channel for the rest of this section: describe, validate, create, publish. in use · builds the rest MSP Ticket Triage It reads the ticket, sorts it, and writes the first reply. The dispatcher still makes every final call. running · private n8n Briefing Line Generating a podcast is one API call now. Refusing to ship a bad one is the part that took engineering. runs weekly · gated Vellum The essay exports as prose. The argument underneath stays queryable. live Alagà · Guest AI Agent It knows the whole stay and it acts with real tools. Anything a guest could get hurt by goes to a person. live run + 4 real apps labor-ph Refusal threshold moved from a hand-picked 0.4 to 0.58 by a 61-step sweep. 0 false refusals at the calibrated line. public · index live, 1,199 chunks Back-office Agent Eight nodes, one interrupt(). 10 of 10 scenarios pass, and a rejection never reaches send. public · gate tested warrant 19 of 19 unauthorized actions reached systems ungoverned. 0 of 19 governed. Same model, same prompt, same tools. public · 19 of 19 to 0 of 19 Landing Page Engine 14 of 14 injected defects caught by the gate, 0 collateral. It also found one check that could not fail. public · gate measured 14/14 meridian-evals gpt-4.1-mini 11 of 11 on every nightly run on record. claude-sonnet-5 misses one or two, and the misses are the story. public · nightly CI
Private production linesFour gated n8n lines on the same rule: the model proposes, code decides, a person approves the send.4 running · private n8n · open ↓ MSP Review & Reputation Engine Routine reviews get fast drafts. The ones that could hurt someone go to a person first. running · private n8n Phishing Sim Production Line Calibrated multi-vector phishing training, produced at volume. running · private n8n MSP SEO Content Line Matrix content with the editor's checklist enforced in code. running · private n8n SprintOS Twelve angles written, eight killed by a critic, four survive to a human. Nothing reaches an ad account. running · private n8n
Tools & productsapps built end to end
gatekeep The two governance properties an AI automation lives or dies on, made into checks a pipeline can block on. live · CLI Reel Point it at a URL, get a clean scroll-walkthrough mp4. One command. live · CLI Dossier The first client call, replaced by a form that produces a brief a designer can start from. live Prompt Maker Prompt guidance that cannot quietly go stale. v1.1 · live Sim Landing Page Generator Looks exactly like the real login. Provably inert. hardened · 3 audits LOIS · Ad Engine Brief in, distinct on-brand Meta batch out, one command. client work · legal-tech Video Ad Line Brief in, four gated video ads out, a human decides, nothing publishes itself. runs on demand · gated
Sitesdesign + front-end
UI/UX case studies The design half of the practice has its own page: 5 sites designed and built end to end, live and clickable. own page →
Body of worksecurity-awareness at scale
Cywareness Simulation Archive Production security-awareness content at volume, and the pipeline that now builds it. 2025 to 2026 · body of work
what i actually run claude as8 capabilities
Agent loop, hand-built

A TypeScript agent on the Anthropic Messages API, no agent framework: the tool-use loop, the memory that spans a whole guest stay, and the guardrail that force-escalates money, safety, and legal before the model gets a turn are all mine. LangGraph is in the Back-office Agent, where a graph earns its keep.

Alagà · Guest AI Agent →
MCP, as a client

Wired to my n8n instance through its official MCP server. A workflow gets described in a sentence, written as SDK code, validated, created, and published without opening the editor.

n8n Workflow Builder Agent →
MCP, as a server

The page engine registers three tools over stdio, so any MCP client can call the same draft-and-grade pipeline I use.

Landing Page Engine →
Skills, orchestrator, agents

Seven Claude Code skills hold the rules, an eight-step orchestrator runs the contract, and twelve sub-agents do the judgment-heavy authoring. The pipeline is public as a sanitized rebuild, with none of the client material.

phishing-sim-line ↗
Evaluation

A golden set of real and adversarial incidents scored against both production brains for gate correctness, cost, and latency, with the scorecard committed to the repo and rerun nightly by CI. Its unit tests caught a live security-floor bug that had survived three deployments in two codebases.

meridian-evals →
Guardrails in code

The pattern under six of these builds: the model proposes, code decides. Severity floors, flag counts, confidence thresholds, and escalation rules live outside the prompt, where they cannot be talked out of.

MSP Ticket Triage →
Prompts as artifacts

Each workflow carries its own locked template, and the prompt tool re-syncs Anthropic's live docs so its guidance cannot quietly go stale.

Prompt Maker →
Model-aware operation

Sonnet's thinking block breaks a naive content[0].text parse, max_tokens is tuned per job, and the sim pipeline has one named seam where a better model swaps in while the deterministic half stays byte for byte.

Cywareness Simulation Archive →