case_file_001 · forward-deployed ai lead

One person.
Two shipped AI products.
No handoffs.

I take ambiguous AI products from decision to deployment — product, architecture, implementation, evals, native release, telemetry, and production outcomes. One product created a category; the other reached the App Store in 14 days.

The two exhibits below are the native App Store releases. Further work — including Yapoleon's Court and the AI Boardroom Forecast Audit — is on the main workfolio.

1st + onlyapp store word game with live-grid ai
14 daysconcept to app store
2,309tests green in captured build
100%ownership, product to production

exhibits a-b / shipped products

Two products. Two hard constraints.

Yapword keeps a generative character inside live gameplay without giving the model control of game truth. That's My Best creates a social-photo product without biometric identity, account access, or staff photo review.

CF-AB-YAP-001live
Y

Yapword

Continuously interactive AI word game

A character inside the deterministic game loop.

The App Store's first and only word game where a generative-AI character reads the live letter grid and responds throughout the entire game.

  • First and only on the App Store
  • AI reads and responds to the live grid
  • 22-rendition voice ledger, R0-R13 roast lab
CF-AB-TMB-002live
TMB

That's My Best

Multimodal AI friend quiz

I removed the dependency instead of negotiating with it.

Creator-supplied Instagram-grid screenshots become a playable friend quiz without face matching, Meta API access, or staff reviewing private photos.

  • No face matching, no Meta API
  • Concept to App Store in 14 days
  • Conservative exclusion of uncertain imagery

01 / one accountable human

Agents multiplied throughput. They never owned a consequence.

AI development agents did not own a decision, a release gate, or a production outcome. Every link in this chain resolved to one person.

01

Product decision

02

Architecture call

03

Implementation

04

Evals + tests

05

Native release

06

Production outcome

exhibit a1 / category creation

In Yapword, the AI stays in the game.

Most AI word games use a model before or after play. Here the rules engine owns the board and the outcome while the character owns the live voice — opponent, narrator, hint system, relationship memory, and postgame critic.

01Live input

Every guess changes the grid

Letters, placements, attempt, mode, standing, and player history become bounded input.

02Engine facts

Truth stays deterministic

Validity, score, board truth, completion, standing, and memory remain outside the model.

03Model output

Voice stays generative

A board-aware reaction, contextual hint, or postgame critique returns without authority over the outcome.

artifact / the voice ledger

The character was engineered, not merely prompted.

A chronological production ledger traces every architecture, register, dial, roast-lab, surface, and incident change beside the real lines each version produced. The public exhibit keeps proprietary prompt text redacted while showing the engineering method and its receipts.

22

renditions logged

R0-R13

controlled roast ladder

14

live production lines

2,309

tests green in captured build

01Direction

Positive mechanisms beat the banlist.

A long enumerated denylist flattened cadence and wit. Replacing it with a positive rule — what comic move to make — restored specificity without surrendering the boundary.

02Evaluation

Same board. Controlled context. Different line.

Identical game states were replayed across voice renditions and six relationship standings, making memory and register changes visible instead of relying on taste and recollection.

03Architecture

The engine computes the world. The model authors the voice.

Deterministic software emits facts and state; the model turns them into language. That boundary keeps personality adaptive without letting the model invent score, memory, or game truth.

exhibit b / constraint-driven architecture

No face matching. No Instagram login. No Meta API. No staff looking at private photos.

Shipped from concept to the App Store in 14 days by replacing facial recognition and Meta access with user-controlled screenshots. The pipeline segments the grid, excludes uncertain imagery, creates the quiz, and bounds what the model's proposed answer key can reach.

Rejected system

Recognize people, connect the account

  • Facial identity matching
  • Instagram OAuth and token custody
  • Meta API and policy dependency
  • Human moderation queue
Shipped system

Understand the screenshot, never the identity

  • User-controlled grid screenshots
  • Automated tile and eligibility pipeline
  • Conservative exclusion of uncertainty
  • Invariant-bounded answer keys
01

User-owned input

The creator supplies 1-4 screenshots from a grid they already control.

02

Tile segmentation

The workflow isolates individual posts and rejects interface chrome, partials, and duplicates.

03

Conservative eligibility

Unsuitable content, apparent minors, and uncertain images are excluded from automated selection.

04

Multimodal ranking

The system favors objective, visually grounded memories with strong quiz potential.

05

Structured generation

Selected images become answerable questions, choices, and contextual reactions.

06

Sealed under invariants

The default path seals the model's proposed key; a paid override lets the creator replace it first. Either way the player payload carries no answers and the seal is immutable — without staff reviewing the photos.

The automation does not need to know who anyone is. The operator does not need to see the user's photos. The platform does not need to grant account access. Safety stays conservative, and the invariants hold whether or not the creator edits the key.

decision record / model governance

The model performs. The deterministic engine governs.

Production AI becomes dependable when authority is explicit. Models generate useful content; they do not own truth, scoring, safety, or spend.

01

Truth

Server-owned and invariant-bounded

The model can propose and perform; code decides the payload, the reveal, and the score.

02

State

Deterministic outside the model

Scores, memory, completion, and product rules remain inspectable.

03

Safety

Untrusted content isolated as data

Prompt-injection boundaries and explicit gates protect control flow.

04

Reliability

Failure paths designed in advance

Fallback chains, classified retries, budgets, and cached degradation.

05

Economics

Cost is a product constraint

Output caps, rate limits, attempt budgets, and measured generation cost.

06

Verification

Release gates across the stack

Contracts, invariants, accessibility, smoke paths, and release gates.

exhibit c / search distribution

Shipping is not enough. The work has to be found.

I treat technical SEO and generative-engine optimization as product infrastructure: make the facts crawlable, the entities legible, the answers citable, and the feedback measurable.

01

Technical SEO

Indexability is engineered

Server-rendered facts, crawl controls, canonicals, sitemaps, structured data, performance, and clean information architecture.

02

GEO

Citation surfaces are designed

Source-qualified claims, entity clarity, extractable answer blocks, crawler policy, and repeated verification across answer engines.

03

Topical systems

Coverage without content sludge

Useful taxonomies, service clusters, multilingual pathways, internal-link graphs, and explicit editorial standards.

04

Measurement

Telemetry closes the loop

Cross-platform events, funnels, session replay, analytics parity checks, and search signals turn discovery into an operating system.

observability loop

Connect model output to what people actually do.

Telemetry is the cross-platform operating record for product behavior — not a pageview counter. Anonymous session, game, prompt, deployment, and relationship identifiers connect AI behavior to completion, abandonment, replay, sharing, registration, and retention.

Unified telemetry

Web + native iOS

Platform-tagged events make game starts, submissions, completions, hints, shares, registration, and purchases comparable across surfaces.

Concrete diagnosis

Replay to interface fix

Session evidence exposed rage clicks on a keyboard that looked live while validation was blocking it; the UI state was then bound to the real validation state.

Measurement contract

Parity, filters, privacy

Production allowlists, analytics parity checks, explicit event contracts, and no-secret / no-PII payload rules keep the numbers decision-grade.

exhibit d / measurement

Three products. Three things that can't be checked automatically.

Every product has one quality property no test can assert — whether the voice is good, whether the scoring is fair, whether it's safe to publish. So I build the grader for that property, and then I take it out of the product. Three products in six weeks; read them in order, because the grader work matures across them.

01

Yapword

started 2026-05-23

Is the character's voice good — and what does it cost?

the grader

A multi-turn bench that walks whole games through the real prompt builder, sweeping the model's thinking level.

the result

~82% of inference COGS was invisible thinking. Zero thinking preserved the voice. The highest setting burned 8,991–11,663 tokens per game — $80–105 per 1k games — and was not better.

02

Yapoleon's Court

started 2026-06-15

Is the scoring rubric fair, learnable, and un-gameable?

the grader

Yapword's harness, forked and re-pointed at a different measurement: 30 demands x weak/mid/strong, plus a fixed-mold probe, against the live judge.

the result

72.2% mid / 0% weak / 96.7% strong — a clean skill gradient. Fixed mold 0% off-axis. 5 of 5 anti-gaming probes pass live. 1,095 calls, 0 errors.

03

That's My Best

started 2026-07-01

Is it safe to publish when the classifier itself is unreliable?

the grader

A fail-closed state machine with quorum voting over an injectable classifier — extracted out of the product entirely.

the result

Shipped as llm-safety-gate: zero-dependency, MIT, every failure path unit-tested without a network.

the arc

Ad hoc → reused → extracted as infrastructure.

The first harness was written for one product's problem. The second product didn't get a new one — the first was forked and aimed at a question it was never built for, which is the first evidence the method transfers. By the third, the grader stopped living inside the product at all and shipped as a standalone library someone else can inject their own classifier into.

What makes a grader trustworthy with no reference implementation

Code can be graded against a reference build. Voice, fairness, and taste cannot. So the scoring function is validated by its properties instead.

01

Learnable

Better input must score better

Strong 96.7% > mid 72.2% > weak 0%. If the ordering breaks, the rubric is broken.

02

Not template-farmable

The negative control

One fixed rhetorical mold, applied to all 30 days, must lose on its off-axis days. It wins 0%.

03

Not gameable

Five adversarial probes, live

Naked flattery, prompt injection, legitimate audacity (the false-positive check), delimiter breakout, grovel-on-economy. All pass.

04

Bounded authority

The model never emits the score

It returns five clamped axis sub-scores; a pure server-side function the model cannot see or reach computes the result.

05

Honest about noise

Distributions, not point values

Hosted inference is not bit-reproducible at low temperature, so every cell runs at least three times and the spread is reported.

The calibration also produced a finding I kept rather than buried: the target was specified as a median win-rate, and the median turned out to be bimodal and degenerate — an offline sweep of hundreds of candidate curves found zero with a median in band. The mean is smooth and stable across independent runs. The document says so, and says why.

exhibit e / distribution surface

A governed agent, and the surfaces it publishes to.

agentic distributionlive

@YapoleonGreater on X ↗

Extends Yapword into a governed, near-autonomous social-publishing system. It generates publication-ready responses in character within owner-defined rules. I retain the voice, the publishing boundaries, the escalation decisions, and the consequences — the same authority split that governs every other model in this file.

Public web properties

Products, tools, and commercial sites — each a real distribution surface. Private systems and retired experiments are intentionally excluded.

prior record / operator judgment

High-stakes ownership came before the code.

More than a decade in commercial real estate taught me to sell complex assets, advise executives, repair operational data, and stay accountable when decisions had real consequences.

Transaction record

85+ retail transactions

Roughly $400M in total volume as half of a two-person team — buying, selling, leasing, and repositioning.

Operating record

5,700+ CRM contacts rebuilt

Roughly 1,900 mapping errors identified and corrected.

Education

Texas A&M, Finance

Commercial judgment and technical execution in the same room.

2012-2025 · CRE operator & advisor2025-now · forward-deployed AI lead

next file / the right hard problem

Put one strong owner on the whole problem.

If you need someone who can turn an ambiguous AI opportunity into a shipped, measured product — without losing the last mile between model and customer — let's talk.