Joseph Stickel

Joseph Stickel

Applied AI Engineer, Production RAG, Agents, Evals & Private LLM Platforms

jobtrue.ai/u/mrjstickel/applied-ai-engineer

Joseph Stickel

Applied AI Engineer, Production RAG, Agents, Evals & Private LLM Platforms

Everett, WA  ·  contact@mrjstickel.com  · 
mrjstickel.com · github.com/j-stickel · linkedin.com/in/mrjstickel

Summary

Built Architecture Zero: a white-label AI platform running four live production instances from one codebase - hybrid RAG (vector + BM25 + reranker), department-scoped knowledge bases with role-based access, and local inference on RTX 5090/3090 hardware. Drove retrieval recall from 40% to a measured 100% (stamped in-app); the harness is now 98 questions, cross-family-judged (answers and grades from different labs) - floors: 98.6% recall, 91.0 tuned answers vs an outside-authored locked holdout (11.3-point gap published), honesty 13/13. The instrument is itself calibrated - 32/32 planted-error suite, human baseline (kappa 0.61), cross-lab re-grade (99.0% agreement): numbers are graded before they're believed. Everything ships with security gating, cost observability, nightly-drilled backups, and adversarially verified fail-closed boundaries. Behind it: 13+ years of enterprise support - when things break at 2am, the monitoring, postmortems, and on-call habits are already built in.

Technical Skills

AI / LLM SystemsRAG pipelines (hybrid vector + BM25 fusion), cross-encoder reranking, retrieval eval harnesses, LLM-as-judge answer evals, agentic tool use, QLoRA / PEFT fine-tuning, multi-provider routing (Claude / GPT / Ollama) with failover, local inference (7B to 30B-class on RTX 5090/3090), ChromaDB, embeddings, prompt-injection defense, LLM-judge calibration / meta-evaluation, AI-assisted engineering (Claude Code: custom skills / hooks / multi-agent workflows)
BackendPython, FastAPI, Node.js, Firebase Admin, REST API design, Stripe, SQL, MySQL, SQLite, Alembic migrations
InfrastructureGCP (Compute + Cloud Monitoring uptime alerting), Vercel, Cloudflare (DNS, tunnels, R2 object storage), Docker Compose, Nginx, Let's Encrypt, CI/CD security gating (gitleaks, pip-audit, bandit, trivy), BetterStack monitoring
FrontendNext.js, React, TypeScript, Tailwind, React Native, Expo
Real-TimeFirebase RTDB, Upstash Redis, Server-Sent Events, distributed locking

Experience

Founder & Platform ArchitectJan 2022 - Present

MetaOne Designs (M1D) - Architecture Zero, TheStatic.tv, JobTrue · Remote

  • Architected Architecture Zero, a private white-label AI platform running four live production instances from one core codebase - hybrid RAG (vector + custom Okapi BM25 + cross-encoder reranker), department-scoped knowledge bases with role-based access, multi-provider LLM routing (Claude / GPT / local Ollama) with failover, and local inference (7B to 30B-class models on RTX 5090/3090 hardware).
  • Drove retrieval recall from 40% to a measured 100% (50/50 retrieval-scored, stamped in-app), then grew the harness to a 98-question cross-family-judged instrument (answers and grades come from structurally different labs) - current floors: 98.6% recall, 91.0 tuned answers vs an outside-authored locked holdout with the 11.3-point overfit gap published, honesty cohort 13/13 - via cross-encoder reranking, markdown-section chunking, a full-corpus BM25 leg, and content-addressed delta ingestion that cut re-index costs from ~470 chunks to ~10 per update.
  • Built Eco Mode, a federated retrieval mesh: instances query each other's knowledge in parallel with per-caller scoped keys and fail-closed access boundaries - a public bot can never reach private departments; verified by adversarial testing as the least-privileged caller.
  • Hardened the fleet with automated CI security gating (gitleaks, pip-audit, bandit, trivy), OWASP Web + LLM Top 10 mapping, per-IP rate limiting behind a proxy-aware real-client-IP boundary, and self-audited postmortems - found and closed a real broken-access-control bug, then published the case study.
  • Shipped TheStatic.tv solo: a production Web3 streaming platform - 160+ authenticated API endpoints, GCP Livestream RTMP/HLS pipeline, a cron-driven billing engine with Redis distributed locks and viewer-scaled burn rates, and an ERC-721 contract on Ethereum mainnet.
  • Shipped JobTrue (jobtrue.ai) solo: an AI job-search SaaS with Stripe billing, LLM fit-scoring with a server-side grounding layer that drops hallucinated resume items before storage, and per-user rate limiting - this resume is served from it.
  • Operate everything built - hosting, CI/CD, monitoring, security review, documentation - plus nightly encrypted backups to immutable object storage (Cloudflare R2 bucket-lock + Google Drive) and an automated daily restore drill that boots each backup in a throwaway container: 9/9 checks, alarmed end-to-end via a fail-closed status probe.
  • Calibrated the LLM-as-judge eval instrument against four independent references - a 32-case planted-error suite (32/32 caught), a human-adjudicated baseline (Cohen's kappa 0.61), a cross-lab re-grade (99.0% agreement, kappa 0.942), and an independently-authored harness (98.2% boundary agreement) - then audited my own locked exam, retired it for authorship bias, re-authored it via a third lab, and published the wider honest gap.
  • Designed refuse-vs-fabricate honesty evals - cohorts demanding artifacts the corpus does not hold, scored by a purpose-built judge rubric with planted calibration - measured 12/12 and 13/13; separately drove an adversarial tier-isolation cohort from 3/11 to 11/11 by locating an access leak diffused across the corpus, with owner recall held at 100%.
  • Rolled the measured trust layer across all four production instances, validating each port with like-for-like A/Bs on that instance's own cohort: +12 correctness on one peer, +18.4 recall with rank-1 hits doubled on another, and multi-turn follow-up retrieval driven 1/9 to 8/9 by a measure-gated resolver port.
  • Eliminated hand-set status from the platform's trust dashboards: every displayed number derives from stored eval runs or live config at request time; 42 reliability-layer claims each self-verify against activation-keyed code proofs on every deploy; the tuned-vs-holdout overfit gap is published as a first-class metric even when negative.
  • Root-caused silent production failures behind green health checks: an HNSW vector-index corruption (search dead while reads passed - shipped a knn tripwire plus per-run corpus fingerprinting) and two distinct vector-store data-loss mechanisms (flush-boundary loss on unclean stops; a file-watcher mass-purge during git deploys) - fixed and verified with back-to-back zero-loss deploy cycles.
ProctorFeb 2026 - Present

Conduent supporting Boeing · Everett, WA

  • Maintain operational availability of testing workstations for aerospace manufacturing candidate assessments with a zero-downtime requirement.
  • Serve as primary on-site contact for all technical issues; resolve hardware, peripheral, and connectivity problems and escalate to remote engineering when needed.
  • Enforce strict site security and technical protocols as the on-the-ground liaison for off-site infrastructure managers.
Technical Support SpecialistSept 2025 - Feb 2026

Decentraland Foundation · Remote

  • Provided technical support for creators and developers building on the Decentraland metaverse platform across Intercom, Discord, and Slack.
  • Troubleshot SDK issues, scene deployment problems, wearables, and platform bugs; documented reproductions and coordinated with engineering teams for resolution.
  • Leveraged 4+ years of hands-on Decentraland development experience to give expert-level guidance to the creator community.
Technical Support Team Lead / Support Engineer IINov 2022 - Jun 2025

Symetrix, Inc. · Mountlake Terrace, WA

  • Drove live call answer rate from under 25% to over 80% by rebuilding NetSuite case-routing workflows and KPI tracking as primary NetSuite administrator.
  • Cut Return Authorization turnaround from several weeks to 1-2 weeks by redesigning the process mid-2024.
  • Hired as Tier 2 support; promoted to Team Lead in June 2023. Recognized publicly by the VP of Sales & Marketing for the call and case routing overhaul.
  • Completed a NetSuite-ZOHO integration proof of concept; 2024 review: Exceeds Expectations across quality, initiative, technical skills, and dependability.
Senior Support EngineerDec 2018 - Mar 2022

SAP Concur · Bellevue, WA

  • Maintained 98% customer satisfaction across a high-volume enterprise Tier 3 queue for SAP Concur expense and travel products.
  • Served as Comms lead on the Incident Response Team for P1/P2 issues - owned customer and stakeholder communication under SLA pressure while Tech and Lead drove resolution.
  • Queried client data with SQL to diagnose enterprise issues during Tier 3 troubleshooting; mentored team members and built performance dashboards.

Selected Projects

Architecture Zero

mrjstickel.com/projects/az

Private white-label AI platform | 4 live instances | measured 100% retrieval recall

  • Engineered hybrid retrieval: vector search over-fetches a candidate pool, a from-scratch Okapi BM25 leg runs across the full corpus, and a local cross-encoder reranks ~60 candidates to the best 5 - recall driven 40% -> 93% -> 100% (50/50 retrieval-scored, stamped in-app); the harness has since grown to 98 questions, cross-family-judged - current floors: 98.6% recall, 91.0 tuned answers, 11.3-point holdout gap published.
  • Built content-addressed delta ingestion: only new or changed text embeds on updates (~10 chunks instead of ~470), a mid-ingest failure resumes cleanly, and every ingest emits an observability event.
  • Designed department-scoped knowledge bases with tiered access and a federated mesh (Eco Mode): parallel peer fan-out, per-caller scoped keys, fail-closed boundaries, graceful degradation when a peer is down.
  • Secured the fleet: OWASP Web + LLM Top 10 mapping, CI gates (gitleaks, pip-audit, bandit, trivy), guest budget caps, per-IP rate limits, injection checks, and adversarial self-audits with published postmortems.
PythonFastAPIChromaDBOllamaONNXDockerNginxGCP

Northwind AI (live demo)

northwind.mrjstickel.com

White-label client demo of Architecture Zero - click it, ask it anything

  • Live public deployment proving the white-label model: a fictional company's branded assistant with real department access boundaries - the "View as" switcher shows Sales seeing pricing docs a public visitor is refused.
  • Public trust panel (northwind.mrjstickel.com/#trust) derives every number from stored eval runs at request time - per-cohort scores never blended, corpus-fingerprinted, and the tuned-vs-holdout overfit gap published even when negative.
Architecture ZeroRAGRBACDockerGCP

Kintsugi (private)

Personal AI over a life's knowledge - the private flagship instance

  • Mobile-first (React Native/Expo) private assistant over a personal corpus: session history, dated-log chunking with recency weighting, department routing, GPU-tunneled local inference (Cloudflare tunnel to RTX 5090), and the eval discipline the case studies document.
  • Private by design - the live public counterpart to walk through is Northwind AI.
PythonFastAPIReact NativeExpoOllamaCloudflare

JobTrue

jobtrue.ai

AI job search platform · Solo build · Stripe billing

  • Solved AI resume hallucination structurally: the LLM selects from the user's original bullets by index and application code reassembles verbatim text - a server-side grounding layer drops and logs any fabricated item before storage.
  • Identified and remediated a critical SSRF vulnerability in the URL scrape route - blocks private IP ranges, loopback, link-local, and internal hostnames before fetch to prevent cloud credential exposure.
  • Shipped Stripe billing (checkout, idempotent webhooks, customer portal), per-user AI rate limiting via Redis sliding windows, fit-score calibration with two-shot anchoring, cost-routed models, and a full admin console with token-cost monitoring.
  • Built the application pipeline tracker: frozen resume + cover letter snapshots per submission, interview rounds, contacts, and fit-score history.
Next.jsFirebaseClaude APIStripeVercel

MySQL Client & SQL Trainer (private)

Full-stack database tool · Solo build · Mobile + API

  • Mobile MySQL client with an AI assistant that reads the actual schema and writes or explains SQL - constrained read-only, Fernet-encrypted credentials, multi-statement blocking, explicit write-mode opt-in.
  • 40-challenge SQL trainer on an isolated in-memory SQLite database, graded by automated row-set comparison.
React NativeExpoFastAPIMySQLSQLiteClaude API

TheStatic.tv

thestatic.tv

Production Streaming Platform · Solo build · 160+ API endpoints

  • Engineered GCP Livestream integration: RTMP ingest, tier-based HLS transcoding (720p-1080p60), CDN delivery, and batched VOD archival for streams of 5,400+ segments.
  • Architected the Spark Economy billing engine: Redis distributed locks (120s TTL), viewer-scaled burn rates (1x to 50x), a cron Governor deducting credits every minute, and Polygon crypto payments with verification windows.
  • Deployed the ZeroPointPass ERC-721 contract to Ethereum mainnet with EIP-2981 royalties; on-chain ownership syncs to platform tier every 12 hours.
  • Unified multi-auth: MetaMask signatures (EIP-191), Decentraland signed-fetch, and Firebase custom tokens in a single authorization layer.
Next.jsFirebaseGCPPolygonRedisEthers.js

What I Deliver

· Production AI systems end-to-end: RAG pipelines with measured retrieval quality (eval-stamped recall + LLM-judged answer floors), agents with tool use, guardrails, and cost observability
· Private / self-hosted AI platforms - your data, your brand, your infrastructure (Architecture Zero: four live instances from one core)
· Retrieval engineering: hybrid vector + BM25 fusion, cross-encoder reranking, section-aware chunking, delta ingestion - fixed at the corpus, not patched at the prompt
· Continuous AI evaluation: question sets, regression floors, measured before/after on every change
· LLM fine-tuning through the full QLoRA lifecycle - and the senior call of when NOT to fine-tune
· Multi-provider LLM architecture: Claude, GPT, and local Ollama with failover; local inference on 24GB-class GPUs
· Federated / multi-tenant AI: department-scoped knowledge, role-based access, fail-closed boundaries - adversarially self-tested
· AI security in production: OWASP LLM Top 10 mapping, prompt-injection defense, CI security gates, published self-audit postmortems
· Documentation as infrastructure: single source of truth, generated facts that cannot drift, docs ingested into the AI itself - measured onboarding drop from 3 weeks to 3 hours
· P1/P2 incident communication and SLA-driven operations - 13+ years of enterprise support instincts wired into everything above
· Ops resilience as machinery: immutable nightly backups, an automated daily restore drill, and fail-closed status probes wired to alerting - an untested backup is a hope

Education

BS, Information Technology Management (BSITM)Jun 2026 - Expected 2027

Western Governors University (WGU)

Returning student completing the degree while building and operating production AI systems full-time.

Built with JobTrue · Create your free resume →