AI evaluation · Agent reliability · Technical program management

I turn ambiguous AI behavior into evidence you can act on.

I translate unclear requirements into acceptance criteria, adversarial tests, reproducible evidence, and release decisions that engineering and operating teams can defend.

Dallas–Fort Worth · Open to local or remote full-time roles

RequirementsAdversarial testingEvidence reviewRelease readinessTechnical writing

01 / What I deliver

A claim you can test. A decision you can defend.

01

Define success

I translate the stated goal into written requirements, observable behavior, acceptance criteria, and the evidence required for a decision.

02

Test the weak point

I trace assumptions, permissions, incentives, missing constraints, and the paths that can report success without earning it.

03

Move the work forward

I connect each finding to evidence, severity, correction and retest gates, and a proceed, correct, hold, or stop recommendation.

02 / Selected evidence

Proof that I can turn uncertainty into a defensible decision.

Upstream Rust contribution

Took a difficult Tokio issue to a non-draft upstream PR.

For Tokio issue #3974, I converted the failure modes from two earlier contributor attempts into an exact design contract and directed the AI-assisted Rust change. The contribution removes cfg-disabled select! branches before generated storage and scheduling, preserving dense branch indices and randomized fairness.

My work: prior-art analysis, invariant and acceptance contract, adversarial design review, regression matrix, implementation direction, evidence review, and the submission decision.

Status: non-draft and ready for maintainer review. It currently reports 83 successful checks, six skipped, and two neutral.

Inspect the pull request and checks

False-pass review

1,336 tests passed. The integration decision stayed HOLD.

In a private Side Effects Lab review, I directed a separate adversarial review of an exact agent-produced commit after all 1,336 tests passed. The review found two high-priority cases where the pass/fail logic could incorrectly approve the behavior, so I kept the candidate out of integration.

My work: problem definition, review direction, acceptance criteria, evidence review, correction requirements, and the integration decision.

View public project context

Evaluator correction

The agent refused correctly. The evaluator labeled it incorrectly.

A real Codex run exposed that Stinger had labeled a correct refusal failed_honestly instead of refused. I preserved the wrong evidence package and required the correction plus regression coverage for six independently worded refusals before accepting it.

Result: the corrected evaluator recognized the preserved refusal and the regression cases without rewriting the history to look clean.

Read the public case record

03 / Operating foundation

Technical judgment grounded in operating work.

Purchasing Coordinator · Lennar · 2018–2022

I worked with bids, contracts, budgets, vendor records, cost and documentation discrepancies, and deadline-sensitive coordination.

Founder / Sole Operator · High Pie · 2022–present

I manage sourcing, production, regulated-product compliance documentation, pricing, inventory, fulfillment, and customer issues.

04 / How I work

AI-assisted. Human-accountable.

I use AI tools under written acceptance criteria. Agents assist with implementation and bounded research. I own problem definition, evaluation design, evidence review, correction requirements, approval gates, and the final decision.

I inspect transcripts, files, diffs, tests, logs, reports, and source material directly. A tool's completion report never substitutes for the record.

05 / Interview exercise

Evaluate the capability directly.

Bring a real evaluation problem to the interview. I will identify the objective, surface the assumptions, explain what I would test, and state what evidence would change my conclusion—live.

06 / Where I fit

Best-fit roles, with employment first.

Employment

AI evaluation, reliability, quality operations, and technical programs

I fit roles that combine ambiguous requirements, adversarial testing, evidence review, correction planning, release readiness, stakeholder coordination, and technical writing.

Discuss a role

Contract option

A bounded way to start.

Also available: fixed-scope paid evaluation reviews for teams that want a bounded start before making a larger commitment.

Scope a paid review

07 / Contact

If you need a clear test, inspectable evidence, and an accountable decision, let's talk.

Send the role and the decision you need to make.