Christine Straub · Lead AI/ML Engineer · Remote / United States

Lead AI/ML Engineer building the evaluation infrastructure that makes AI agents trustworthy in production.

I design RL environments, deterministic verifiers, and benchmark tasks that measure whether AI agents and LLMs actually work, and I ship the production systems behind them: document intelligence pipelines (OCR, VLM extraction), RAG, MCP-based AI agents, GPU-optimized inference, and MLOps/LLMOps.

Open to remote AI/ML Engineering, AI Agent, and ML Evaluation roles

  • Document Intelligence
  • VLM / OCR
  • AI Agents + MCP
  • RL Environments
  • RAG + Evaluation
  • vLLM / GPU Inference
  • MLOps / LLMOps
  • Python / PyTorch / GCP / AWS

Selected outcomes

years building production AI/ML and software systems
0+
PRs reviewed across production systems
0+
production bugs resolved
0+
images/day supported in ML pipelines
0M+
deployment-time reduction through MLOps automation
0%
infrastructure cost reduction in serverless ML/API architecture
0%
extraction accuracy target in document intelligence workflows
0%+

Case Studies

Production systems, evaluated.

Evaluation infrastructure first, then the agentic, document, and MLOps systems it protects. Every case includes measurable outcomes and a technical deep dive.

RL Environments & AI Evaluation Infrastructure

Bespokelabs AI · Micro1 · Handshake AI

Designed realistic RL environments, task specifications, reward signals, verifiers, hidden tests, and benchmark tasks for AI coding agents and model evaluation workflows.

  • Python
  • PyTorch
  • JAX
  • Hugging Face
  • verifiers
  • Docker
  • pytest
  • CI
  • benchmark harnesses

Outcomes

  • Authored realistic multi-step coding-agent repair tasks
  • Built deterministic verification and golden-reference checks
  • Analyzed rollouts for failure modes and reward hacking
  • Created evaluation criteria for reproducibility, correctness, and robustness

AI Agents, MCP Tooling, and Multi-Agent Orchestration

Unstructured IO · Bespokelabs AI · Handshake AI

Built and evaluated agentic systems using MCP servers, Pydantic AI, tool-calling workflows, deterministic verification, agent rollouts, and failure-analysis loops.

  • Python
  • TypeScript
  • MCP
  • Pydantic AI
  • FastAPI
  • Docker
  • pytest
  • OpenAI
  • Claude
  • Gemini

Outcomes

  • Designed MCP-based workflows for tool use and controlled agent execution
  • Evaluated agent failures including wrong tool selection, bad arguments, and reward hacking
  • Created verifier functions, test harnesses, and reproducible benchmark tasks
  • Improved agent reliability through structured evaluation and debugging

Document Intelligence & Multimodal Extraction

Medici Land Governance · Unstructured IO

Built document AI systems for deeds, liens, mortgages, court dockets, PDFs, and complex enterprise documents using Gemini Document Intelligence, Qwen VLMs, OCR pipelines, layout detection, schema-constrained extraction, and human-review workflows.

  • Gemini
  • Vertex AI
  • Qwen3-VL
  • vLLM
  • Instructor
  • Pydantic
  • OCR
  • FastAPI
  • Python
  • GCP

Outcomes

  • Designed hybrid OCR → VLM → frontier-model extraction pipelines
  • Used confidence scoring and validation checks for reviewable outputs
  • Supported structured extraction for high-stakes legal and enterprise documents
  • Improved reliability through schema validation and benchmark-driven evaluation
  • Contributed open-source OCR wrappers (PaddleOCR, Tesseract) and parsing pipelines across seven Unstructured repositories

Production MLOps, Inference, and Computer Vision Systems

RIOS Intelligent Machines · Sapient Logic · Collegis Education

Built scalable ML systems for computer vision, OCR, speech analytics, cloud ETL, real-time inference, and deployment automation.

  • Python
  • PyTorch
  • YOLO
  • ONNX
  • TensorRT
  • Kubernetes
  • Metaflow
  • GCP
  • AWS
  • Docker

Outcomes

  • Reduced deployment time by 70%
  • Supported real-time 60 FPS ML inference workflows
  • Processed 10M+ images/day
  • Reduced manual annotation by 60% through active learning workflows
  • Optimized models for GPU and edge deployment

Experience

Timeline.

Nine roles across defense, healthcare, finance, robotics, legal document workflows, and enterprise AI platforms.

  1. Lead AI/ML Engineer · Medici Land Governance

    Document intelligence for land records and legal documents

    • Built document AI for deeds, liens, mortgages, and court dockets
    • Designed hybrid OCR → VLM extraction with confidence-scored review
    • Improved reliability through schema validation and benchmark-driven evaluation
    • Gemini
    • Qwen3-VL
    • vLLM
    • Pydantic
    • FastAPI
    • GCP
  2. Lead AI/ML Engineer · Bespokelabs AI

    RL environments and evaluation infrastructure for coding agents

    • Designed realistic RL environments, task specs, reward signals, and verifiers
    • Analyzed agent rollouts for failure modes and reward hacking
    • Created evaluation criteria for reproducibility, correctness, and robustness
    • Python
    • PyTorch
    • JAX
    • verifiers
    • Docker
    • CI
  3. Senior ML Engineer · RIOS Intelligent Machines

    Computer vision and robotics for industrial automation

    • Built scalable ML systems for computer vision and real-time inference
    • Reduced deployment time by 70% through MLOps automation
    • Processed 10M+ images/day and optimized models for GPU and edge deployment
    • PyTorch
    • YOLO
    • TensorRT
    • ONNX
    • Kubernetes
    • Metaflow
  4. Senior AI/ML Engineer · Unstructured IO

    Enterprise document processing and agentic workflows

    • Built multimodal extraction pipelines for complex enterprise documents
    • Designed MCP-based workflows for tool use and controlled agent execution
    • Created verifier functions, test harnesses, and reproducible benchmark tasks
    • MCP
    • Pydantic AI
    • OpenAI
    • Claude
    • FastAPI
    • pytest
  5. Lead SWE / AI/ML · Sapient Logic

    ML platform engineering and cloud ETL

    • Built cloud ETL pipelines and deployment automation for production ML
    • Reduced infrastructure cost by 80% in serverless ML/API architecture
    • Optimized models for GPU and edge deployment
    • Python
    • GCP
    • AWS
    • Docker
    • serverless
  6. AI Software Architect · Speechlab AI

    Speech analytics and audio intelligence

    • Built scalable ML systems for speech analytics
    • Supported real-time inference workflows in production
    • Improved reliability through structured evaluation and debugging
    • Python
    • FastAPI
    • Docker
    • GCP
  7. ML Engineer · Collegis Education

    Applied machine learning for education platforms

    • Built production ML pipelines for cloud ETL and inference
    • Reduced manual annotation by 60% through active learning workflows
    • Reduced deployment time through MLOps automation
    • Python
    • PyTorch
    • Metaflow
    • AWS
  8. Software Engineer · Moody's Analytics

    Financial data platforms and analytics software

    • Built production software systems for financial analytics
    • Reviewed 500+ PRs across production systems
    • Resolved 300+ production bugs
    • Python
    • PostgreSQL
    • APIs
    • distributed systems

Skills

Technical matrix.

The tooling behind the systems above, grouped by how it is used in production.

AI Engineering

  • LLMs
  • VLMs
  • RAG
  • embeddings
  • reranking
  • prompt engineering
  • context engineering
  • AI agents
  • MCP
  • tool calling
  • evaluation
  • RL environments
  • verifiers
  • LangGraph
  • LlamaIndex
  • DSPy

Document Intelligence

  • OCR
  • layout detection
  • structured extraction
  • schema validation
  • confidence scoring
  • human-in-the-loop review
  • PDF pipelines
  • Tesseract
  • PaddleOCR

ML / Deep Learning

  • PyTorch
  • JAX
  • Hugging Face
  • fine-tuning
  • LoRA/PEFT
  • computer vision
  • NLP
  • active learning
  • TensorFlow

Inference / Optimization

  • vLLM
  • TensorRT
  • ONNX
  • quantization
  • A100/H100
  • batching
  • KV-cache
  • latency/throughput benchmarking

MLOps / LLMOps

  • Docker
  • Kubernetes
  • CI/CD
  • Metaflow
  • monitoring
  • evals
  • observability
  • regression testing
  • MLflow
  • Weights & Biases
  • LangSmith

Backend / Cloud

  • Python
  • FastAPI
  • TypeScript
  • Node.js
  • PostgreSQL
  • MongoDB
  • GCP
  • AWS
  • serverless
  • APIs
  • event-driven systems

Proof of Work

Open source & artifacts.

Evaluation infrastructure, environments, and pipelines built in the open, with merged contributions across the Unstructured ecosystem.

Open source: Unstructured

7 repositories · dozens of merged PRs

Substantial contributions across seven repositories in the Unstructured ecosystem (the leading open-source toolkit for turning unstructured documents into clean, structured data), spanning the full intelligent document processing pipeline, from OCR and layout modeling to API design and SDK tooling.

  • Built and optimized parsing pipelines for PDFs, images, emails (EML), and HTML, with layout-aware extraction for downstream NLP
  • Developed and maintained the PaddleOCR and Tesseract wrappers (unstructured.PaddleOCR, unstructured.pytesseract) and backend-agnostic OCR abstractions
  • Contributed to the unstructured API and unstructured-js-client: RESTful parsing endpoints, consistent element-metadata formats, improved async workflows
  • Maintained Docker base images for reproducible, production-ready OCR and inference deployments

Dozens of merged pull requests across seven repositories, plus issue triage, code reviews, and CI improvements.

Benchmark task authoring

Versioned benchmark tasks with deterministic verifiers and golden references for evaluating coding agents.

RL environment design

Realistic multi-step RL environments with reward signals, hidden tests, and failure-mode analysis loops.

MCP tools and agent workflows

Typed, auditable MCP servers and tool-calling workflows for controlled, inspectable agent execution.

Document AI / VLM evaluation

Benchmark-driven extraction evaluation with schema validation, confidence scoring, and regression checks.

Production ML pipelines

Reproducible training and deployment pipelines with CI, monitoring, and measured inference optimization.

Résumé

Full employment history and credentials, one page.

Download PDF

Project Archive

Nine years of shipped systems.

40 additional engagements behind the case studies above: the full depth, grouped by domain, in scannable form.

AI, ML & document intelligence

14
  • Unstructured

    Open-source document intelligence: parsing, OCR, layout inference, table extraction, API reliability, and SDK work across 7 repos.

  • Medici Land Governance

    Model-routing document pipelines for land records: OCR, VLMs, schema enforcement, and cost/accuracy escalation (Vertex AI, Gemini, Qwen3-VL, LoRA).

  • Sapient Logic

    Cyber-defense NLP: semantic mapping of threat descriptions to MITRE ATT&CK, with embeddings, ranking, and mission OCR for sensitive environments.

  • PlusOne

    Speech analytics at call-center scale: ASR, keyphrase extraction, sentiment, BERT modeling, and BigQuery ETL on thousands of calls.

  • DQLabs / Intellectyx

    Semantic column-type detection for enterprise tabular data: Sherlock/Sato-style models that classify fields from data content, not headers.

  • SpeechLab

    Event-driven serverless gateway for ASR, machine translation, and TTS on AWS, with cost and performance optimization.

  • EMCA / Energy Log Server

    Time-series forecasting over Elasticsearch operational metrics: LSTM, ARIMA, and anomaly/prediction workflows.

  • Specific Diagnostics

    Medical sensor-image analysis: converting raw imagery into colorimetric time-series with quality-control detection (TensorFlow, OpenCV).

  • Vasoactive Image Analysis

    Microvascular before/after analysis: image registration, vessel skeletonization, diameter-change measurement, and PDF reporting.

  • AI Equity Research Analyst

    Streamlit LLM app that reads 10-K filings via LlamaIndex sub-question retrieval and assembles structured equity-research sections.

  • Agentic Slides / Duarte AI

    Agentic slide-generation and research automation prototypes with guardrails-oriented workflows, producing PPTX, text, and image artifacts.

  • Understanding Patient Conversation

    Rasa conversational AI for medical questions: intent/entity recognition, chief-complaint matching, and dialogue management.

  • AI Customer Assistance

    OCR/NLP system for building-defect communication: extracting and classifying specification documents into assistance workflows.

  • SHM Foundation

    NLP email classification for nonprofit operations, with text-dataset preparation and requirements-driven ML project structure.

Computer vision & edge AI

8
  • RIOS Intelligent Machines

    Industrial robotics vision and MLOps: real-time YOLOv8/v9 inference, TensorRT/ONNX edge optimization, active-learning loops, Jetson deployment.

  • Traffic Monitoring System

    Real-time transportation CV: vehicle tracking, ALPR, lane/crossline detection, and incident reporting (YOLO + Deep SORT, GPU-accelerated).

  • Smart Retail Visitor Flow

    Retail people-counting and analytics: face recognition, re-identification, age/gender detection, and heatmap reporting.

  • Robotic Vision for Farm Work

    Edge robotics: custom detection models connected to motor-control and PID logic, deployed on Jetson, Raspberry Pi, and Odroid.

  • Car Inspection System

    Road-camera undercarriage inspection: video stabilization, optical flow, image stitching, and mosaic vehicle reconstruction.

  • Drone AI for Pineapple Farming

    Drone imagery for precision agriculture: flowering detection, crop counting, and CNN object detection with geospatial analytics.

  • Satellite Image Segmentation

    Landsat urban segmentation from classical ML to Mask R-CNN, exporting GeoTIFF outputs for GIS downstream analysis.

  • Claimant Face Recognition

    Biometric identity workflows combining webcam face recognition, voice recording, feature extraction, and speech-to-text.

Data platforms & engineering

7
  • Collegis Education

    Near-real-time education operations pipelines: BigQuery, dbt lineage, ThoughtSpot reporting, and call-sentiment speech analytics.

  • Houston David Wayne Hooks Airline

    BigQuery data warehouse integrating SQL Server, FTP, Salesforce, and APIs: layered design, CDC/SCD patterns, monitoring and alerts.

  • ONE / Exadata → BigQuery POC

    Oracle Exadata migration proof of concept on Dataflow/Beam, plus a TensorFlow forecasting model for empty-return-yard prediction.

  • Inxeption

    Logistics price-estimation tools plus Python/AWS/Spark ingestion and analytics across S3, Athena, Kinesis, Snowflake, and Redshift.

  • BitReelCo

    Backend and data infrastructure for a 3D showroom and commerce platform: Airflow, API services, SQL migrations, Docker.

  • DeepChannel

    Containerized Python data-processing and prediction workflows with dbt-oriented analytics components.

  • Memetica

    Data and reporting components for digital-investigation and threat-intelligence workflows: Python services, Jinja reporting, Ansible.

Full-stack & product

11
  • MyRuck AI

    Veterans-benefits product in Next.js 14: benefits data ingestion, a Puppeteer scraper converting Army library pages into structured JSON, Prisma/Vercel stack.

  • Duxre

    Multi-application CRE platform: web dashboards, Unity/3D components, APIs, and a Python AI-agent service for asset analysis.

  • COMET

    Electron desktop application for mission workflows: frontend architecture, local packaging, and interface behavior.

  • POLAR

    Mobile capture and translation app: Android/Kotlin CameraX workflows connected to OCR/translation backend logic.

  • AddyCar

    Driver/advertiser marketplace: dashboards, chat, heatmap visualization, and backend services across MEAN and ASP.NET.

  • Cirrent

    React/Redux interfaces with permission-aware behavior, MongoDB data access, and D3 visualizations for network infrastructure data.

  • Perceptyx

    RFP proposal generator: template selection from a database with automated PDF, DOC, and Excel output.

  • Prime Target

    REST services and React visualizations for a global market-intelligence platform (Django/Flask, PostgreSQL, AWS).

  • Oregrown / rVibe

    First-version MERN MVPs with Google Auth, Firebase, and Auth0: full-stack architecture from database to auth flows.

  • CFM / ReturnQueen

    Dockerized full-stack apps across Node/React and Python/Flask: service behavior, state management, and automation bots.

  • Songstream

    First version of a single-page music application with playback, search, playlists, favorites, and queues.

About

About

I'm Christine Straub, a Lead AI/ML Engineer specializing in AI agent evaluation and production machine learning systems. My core work is evaluation infrastructure for AI agents: RL environments, task specifications, deterministic verifiers, and benchmark tasks. I also build the production systems those evaluations protect: document intelligence, OCR/VLM extraction pipelines, RAG, AI agents and MCP tooling, computer vision, and MLOps/LLMOps. Over 9+ years I have shipped production AI/ML systems across defense, healthcare, finance, robotics, legal document workflows, and enterprise AI platforms, with open-source contributions to Unstructured, the leading open-source document intelligence toolkit.

AI systems should do more than demo well. I build for clear inputs, reliable tools, measurable behavior, strong evaluation, and graceful human review when uncertainty is high. If your team needs an AI/ML engineer for agent evaluation, document intelligence, or production ML systems, let's talk.

Lead AI/ML Engineer · Remote / United States

Education & Credentials

Foundations.

Formal training behind the systems above.

Education

  • B.A. in Computer Science

    UC Berkeley

  • B.A. in Cognitive Science

    UC Berkeley

Certifications

  • Machine Learning SpecializationStanford University
  • Deep Learning SpecializationDeepLearning.AI
  • IBM Data Science SpecializationIBM
  • AWS Cloud Practitioner EssentialsAWS
  • Certified Scrum Product Owner (CSPO)Scrum Alliance
  • Google Business Intelligence SpecializationGoogle
  • Google Data Analytics SpecializationGoogle

Contact

Building AI agents, document intelligence systems, evaluation infrastructure, or production ML workflows? Let’s talk.

(949) 527-5247christinemstraub@gmail.com