Staff Software Engineer, Agents Platform at Together AI
San Francisco, CA, United States
I build production AI agent platforms. At Together AI, I established data engineering and created the Agents Platform team, building agent harnesses, tracing, replay, and evaluation infrastructure.
My work connects AI agents with the distributed systems they depend on. I care about reliable execution, useful evaluation, and tools that make engineering teams more effective.
Outside work, I explore the mountains with my dog, build computers, and experiment with new tools.
Experience
Staff Software Engineer, Agents Platform, Together AI
– Present · San Francisco, CA
Together AI is an AI acceleration cloud offering fast inference, fine-tuning, and dedicated GPU clusters for open-source and custom models, backed by a research team that builds and open-sources frontier models and systems optimizations.
Tech lead for data engineering since joining: grew the team from just myself to 15+ engineers while setting technical direction for the data platform behind Together's AI acceleration cloud.
Created the agents platform team after getting the data engineering team off the ground; built the agent harness and internal agents factory to deliver production agents, and contributed to multiple patent-pending innovations in agent systems.
Built and operate agents for infrastructure automation, finance automation and other internal workflows that augment the teams they serve, extending team capacity while improving on-call response and infrastructure maintenance.
Build the data platform behind Together's AI acceleration cloud: the pipelines, storage, and telemetry that turn inference and training traffic into product, reliability, and capacity signals.
Build tracing, replay and evaluation harnesses so internal teams can measure and regression-test LLM agents before they ship.
Build tracing, replay, and regression-testing harnesses for LLM agents, giving internal teams a repeatable way to evaluate agentic features before they ship.
Own core data platform services for inference, fine-tuning, and dedicated GPU cluster workloads, making platform usage measurable end to end.
Develop streaming and batch pipelines that make model-serving and GPU-fleet telemetry queryable in near real time for capacity planning and reliability engineering.
Design data infrastructure for large-scale training and inference workloads, covering dataset curation, lineage, and quality controls for open-model work.
Partner with research, infrastructure, and product teams to standardize data contracts and lineage for training-data and evaluation workflows.
Mentor engineers and set technical direction for the data platform and agents platform architecture.
Prove is the modern platform for consumer identity verification. We enable businesses to securely verify customer identities in real-time, with the highest level of compliance and user experience.
Built agents and harnesses that govern the statistical models used for identity and fraud resolution; these became core frameworks.
Built AI-driven Retrieval-Augmented Generation (RAG) chatbots with Airflow, LangChain, and OpenAI, automating 150+ human-hours weekly.
Engineered AI-powered RAG chatbots leveraging Airflow, LangChain, and OpenAI APIs to automate customer service workflows, reducing manual efforts by 150 hours per week. Addressed NLP challenges in ambiguous query handling, driving a 20% improvement in first-contact resolution.
Spearheaded a company-wide migration from legacy Java/Oracle infrastructure to a cloud-native Go/PostgreSQL stack, cutting API response times from 30s to 12ms and slashing operational costs by 95%. Overcame challenges with service downtime, seamlessly integrating AWS services (Athena, S3, EC2) for enhanced scalability.
Managed and mentored 9 data engineers and 6 data scientists, driving a 20% improvement in project delivery timelines. Collaborated with the VP of Platform Engineering on critical initiatives, serving as the primary data liaison to ensure seamless cross-departmental coordination.
Deployed an event-driven data streaming platform with Go, Kafka, and Flink, transitioning 1,200+ batch jobs to real-time processing. Reduced data availability lag from days to seconds, significantly boosting dashboard accuracy and enabling data-driven decision-making for executive teams.
Implemented real-time telemetry services with Prometheus, Grafana, Splunk, and AWS CloudWatch, addressing complex monitoring needs. Reduced incident response times by 40% while maintaining 99.9% system uptime for mission-critical services.
Designed and developed data pipelines using Spark, Airflow, and DBT to streamline ETL processes, improving performance by 35%. Reduced report generation timelines from days to hours, enhancing reporting capabilities for product managers and business teams.
Established robust data governance frameworks with Apache Atlas, ensuring full GDPR and SOC2 compliance. Implemented metadata standards, lineage tracking, and access policies, reducing audit preparation times by 30%.
Optimized CI/CD pipelines for Apache Beam and Kubernetes deployments, resolving bottlenecks and reducing deployment times by 30%. Improved release frequency to support faster feature rollouts and iterative development cycles.
Migrated computationally intensive SQL queries from RDS to Spark DataFrames, improving query performance by 70% and reducing compute costs by 30%. Enabled seamless processing of large datasets for business-critical applications.
Made up for lost time during COVID-19 by exploring the world, chasing outdoor adventures, and spending time with family.
Made up for lost time during COVID-19 with a road trip through Seattle, Portland, San Francisco, Austin, New York, Washington, DC, Chicago, Los Angeles, and San Diego.
Spent 50 nights camping and hiking, and skied 100 days that season in Colorado and Utah.
Traveled through Europe, spent time with family members in need, and relocated back to Denver from California.
Meta Platforms, Inc. is an American multinational technology conglomerate based in Menlo Park, California. It was founded by Mark Zuckerberg, along with his college roommates and fellow Harvard University students Eduardo Saverin, Andrew McCollum, Dustin Moskovitz and Chris Hughes, originally as TheFacebook.com—today's Facebook, a popular global social networking service.
Built AI/ML classifiers and the data, feature and evaluation pipelines around them.
Served as liaison between Facebook Messenger, Groups, and an early precursor to the Llama project, fine-tuning models for automated group management.
Designed and deployed 100+ ML pipelines using Airflow, Spark, and PyTorch, integrating NLP and computer vision models. Improved ad targeting precision and personalized notifications, driving higher engagement across billions of users.
Lead data engineer for Facebook Public Groups, Community Chats, and cross-platform initiatives, leading 10+ engineers. Collaborated with data science, ML, and hardware teams to align product goals, resulting in a 20% faster delivery of cross-platform features.
Designed exabyte scale data models for Community Messenger to handle multi-platform data streams, increasing user engagement by 35% across Instagram, WhatsApp, Facebook, and Quest despite complex cross-platform dependencies.
Developed modular frameworks for data pipelines, streamlining data flows for thousands of engineers. Reduced integration issues by 40% and accelerated feature deployments by 25% through automation and standardization.
Built telemetry systems capable of processing 10M+ events/second, improving signal quality by 20%. Developed Jinja-based monitoring tools, reducing downtime by 15% and ensuring reliable system performance.
Implemented graph and entity models to support 10 billion monthly interactions, ensuring seamless experiences across Instagram, WhatsApp, Facebook, and Quest. Addressed latency and consistency challenges to maintain real-time performance.
Served as the liaison between Facebook Messenger, Groups, and the early precursor to the LLaMA project, fine-tuning models for automated group management. Increased group engagement by 15% in pilot testing, paving the way for future LLM-powered features.
Created from scratch QR code group invites, increasing join rates by 50%. Enabled offline engagement in multilingual regions, facilitating family reconnections and shelter logistics coordination during humanitarian efforts. Featured on Tech Crunch.
Designed KPI dashboards to monitor DAUs, MAUs, and engagement trends using tools such as Tableau and internal data visualization frameworks. Enabled real-time insights, increasing engagement by 10% and retention by 12%.
Engineered an automated framework for generating thousands of asynchronous Spark data pipelines, increasing compute efficiency by 66%. Overcame orchestration challenges, improving resource utilization and processing times.
Deloitte Touche Tohmatsu Limited, commonly referred to as Deloitte, is a multinational professional services network. Deloitte is one of the Big Four accounting organizations and the largest professional services network in the world by revenue and number of professionals, with headquarters in London, United Kingdom
Automated approximately 31% of human processed healthcare claims using transformer machine learning models, saving 250,000+ hours annually.
Optimized exabyte-scale video reliability metrics, reducing daily processing time by 90% while expanding metric coverage.
Automated 31% of healthcare claims processing with transformer-based ML models, addressing data inconsistencies and regulatory constraints. Saved 250,000+ human-hours annually and improved claim accuracy by 20%
Optimized exabyte-scale video reliability metrics using distributed data processing frameworks. Cut daily processing time by 90% and expanded metric coverage across multiple product lines, improving monitoring precision.
Led teams of 20+ consultants on high-profile Fortune 50 engagements, delivering advanced technical solutions aligned with client needs. Achieved 100% on-time project delivery across multiple engagements.
Developed a risk detection service using machine learning algorithms to flag high-value insurance accounts. Reduced billing errors by 30% and improved revenue recovery through early detection of anomalies.
Engineered NLP models inspired by Google's Transformer architecture within six months of its release. Reduced billing errors by 20% by implementing cutting-edge language models for document processing.
Implemented massively parallel data pipelines using Spark and asynchronous frameworks, reducing healthcare claim turnaround times by 40%. Improved processing efficiency for high-volume workloads.
Re-architected live video infrastructure to align with emerging short-form content trends, saving $100 million and 14 months of development time. Ensured seamless adoption of new formats across the platform.
Optimized caching strategies, leveraging Redis and CDN-layer optimizations to cut costs by $10 million annually. Enhanced content delivery efficiency and reduced latency for high-traffic web services.
Built predictive analytics models for server uptime using time-series forecasting techniques, improving video delivery reliability by 30% and enhancing user experience through proactive maintenance.
Revamped A/B testing frameworks for billions of daily users, resolving data inconsistencies and enabling more precise feature rollouts. Accelerated data-driven decision-making with improved statistical significance tracking.
Leveraged Python, SQL, Java, TensorFlow, PyTorch, Spark, Airflow, Docker, and Kubernetes to develop and deploy scalable technical solutions. Delivered projects across machine learning, real-time analytics, and data pipeline automation for Fortune 50 clients.
Technologies: Python, SQL, JavaScript, TypeScript, PHP, Rust, AWS, Google Cloud
Wide Open West is the sixth largest cable operator in the United States. The company offers landline telephone, Cable Television, and broadband Internet services
Built Machine Learning applications using custom classification and churn models, driving a 22% YoY increase in customer package upgrades.
Led a team of 5 data practitioners, providing BI and data insights to sales, product, and engineering teams company-wide.
Built machine learning models, including custom classification and churn prediction algorithms (e.g., logistic regression, K-means clustering). Increased customer package upgrades by 22% YoY and reduced churn by 15%.
Managed a team of 5 data practitioners, providing BI and insights to sales, product, and engineering teams. Established KPIs and standardized reporting through governance committees, driving a 10% increase in sales performance.
Developed a Kafka-powered real-time analytics platform, streaming data from field technicians and delivering instant job updates via a custom web portal. Reduced service completion times by 40%, replacing 20-minute phone calls with real-time notifications.
Built scalable, cloud-based data solutions leveraging AWS services (SageMaker, S3, Redshift, and Athena). Improved data accessibility and reduced query times by 30% across sales and operations teams.
Automated marketing campaigns using SendGrid to engage at-risk customers, reducing churn by 15%. Implemented multi-channel unsubscribe mechanisms, ensuring 100% compliance with communication preferences.
Delivered geospatial insights using GIS tools to support sales in Arkansas and Alabama. Optimized resource allocation down to the city block level, driving an 18% increase in sales conversions and empowering door-to-door teams.
Developed a dynamic revenue forecasting tool using Python and SQL to set bonus targets and calculate commissions by territory. Increased sales velocity and retention, driving a 12% increase in quarterly sales.
Created a sales funnel dashboard with Tableau to monitor add-on targets, conversions, and installations. Identified millions in unrealized losses, leading to strategic reallocations and improved market performance.
Applied machine learning models and geospatial analytics to optimize network NUC performance, reducing infrastructure build-out costs by 25%. Implemented continuous deployment with GitHub, ensuring code quality through design principles and best practices.
CommonSpirit Health is a nonprofit, Catholic health system dedicated to advancing health for all people. It was created in February 2019 through the alignment of Catholic Health Initiatives and Dignity Health. CommonSpirit Health is the largest nonprofit health system in the U.S. with more than 1,000 care sites in 21 states.
Architected and delivered new rest APIs and data lakes, improving data processing time for external partner data products from 7 days to 5 minutes.
Developed machine learning pipelines using Python and SQL to forecast hospital procedures, billing, and staffing. Reduced billing turnaround from 14 to 7 days.
Architected and delivered new REST APIs and data lakes, reducing external partner data processing time from 7 days to 5 minutes. Enhanced data accessibility and scalability through efficient data structures.
Managed tier-one vendor data extracts containing patient records, financial data, and ICD codes. Improved extract performance by 40% and ensured 100% HIPAA compliance through encryption and access control measures.
Developed a data quality dashboard using Tableau and Python, enabling real-time failure detection. Reduced issue resolution time by 50% and improved overall data integrity and operational reliability.
Conducted advanced data analysis using Dimensional Fact Models in SMP and MPP environments, improving query performance by 35%. Delivered actionable insights to executives, enhancing decision-making processes.
Developed flexible big data extracts and real-time CDC lakes using AWS Redshift and Kafka. Enabled faster product delivery, reducing time to market by 20%.
Designed complex data models for highly sensitive UII data, employing encryption and role-based access control. Ensured data security while supporting high-stakes analytics and compliance use cases.
Created analytics dashboards using Qlik Sense, Tableau, and Python to track key metrics. Delivered actionable insights to executives, increasing reporting efficiency by 25% and enhancing operational visibility.
Developed optimized storage and compute solutions, reducing third-party vendor data costs by 45%. Earned recognition from the VP of Business Intelligence for faster issue resolution and improved efficiency.
Migrated data from relational to columnar formats (e.g., Parquet) using Redshift, improving query speed by 40% and enabling large-scale data processing for analytics.
Developed machine learning pipelines using Python and SQL to forecast hospital procedures, billing, and staffing. Reduced billing turnaround from 14 to 5 days, optimized surgical room usage by 3000%, and enabled predictive staffing to improve patient care.
Worked with orchestration tools similar to Airflow and utilized cloud technologies (Microsoft Azure and AWS) for data infrastructure, achieving seamless cloud operations and reducing deployment times.
Leveraged strong Python and SQL expertise for ETL/ELT processes, developing scalable solutions with continuous improvement and adhering to best coding practices.
R1 RCM is an American healthcare revenue cycle management company servicing hospitals, health systems and physician groups across the United States. Headquartered in Chicago, Illinois, R1 RCM is publicly traded on the NASDAQ.
Built a custom invoicing system leveraging rule-based algorithms and machine learning, driving over $300M in annual revenue.
Migrated clients from SFTP to real-time APIs with Flask and Apache Kafka, reducing data delivery times by 50%.
Developed a custom invoicing system using Python and rule-based algorithms with ML components, generating $300M+ in annual revenue. Improved invoice accuracy by 35% and automated processes to reduce manual effort by 60%.
Partnered with executives on high-impact initiatives, driving a 200% increase in revenue and boosting client retention by 25%. Optimized operational strategies, improving profit margins from 41% to 78%.
Collaborated with CFOs and revenue cycle directors at leading healthcare systems to implement automated reconciliation processes. Cut reconciliation times by 50% and achieved 98% client satisfaction through scalable data solutions.
Built a financial reconciliation platform with Django and Celery, automating line-item invoicing and scheduling. Improved billing speed by 40% and streamlined operations despite complex client requirements.
Led infrastructure migration to AWS, moving PostgreSQL to RDS and compute workloads to EC2. Reduced query latency by 30% and cut infrastructure costs by 25%. Implemented IAM-based security, improving compliance audit performance by 20%.
Refactored 100+ code modules into optimized Python, leveraging multithreading, hashing, and compression techniques. Boosted system efficiency by 45%, resolving bottlenecks in data processing workflows.
Diagnosed and resolved issues in legacy Java applications, reducing downtime incidents by 15%. Enhanced UI responsiveness by 20% through performance optimizations in JavaScript.
Created a Django-powered cron job system with Celery for task orchestration and real-time tracking. Increased task completion rates by 30% through automated monitoring and recovery mechanisms.
Migrated clients from SFTP to real-time APIs with Flask and Apache Kafka, reducing data delivery times by 50%. Enhanced data accessibility, streamlining client operations and improving service quality.
Developed fault-tolerant systems with Spring Boot, implementing state consistency mechanisms such as transaction rollbacks. Reduced system fault impact by 35%, ensuring high availability during critical operations.
Supported pre-sales efforts and customer onboarding with on-site integrations, accelerating implementation timelines by 20% and enhancing customer satisfaction.
A recruiter-facing agent with public MCP and GraphQL context, bounded tools, and reviewed contact actions.
Problem: make professional experience useful to recruiters and their agents. Approach: one approved content source feeds the website, résumé, MCP, GraphQL and grounded agent tools. Tradeoff: keep the public agent narrowly scoped, with signed access, durable usage budgets and reviewed contact actions instead of general execution. Verification: automated backend, frontend and browser checks plus deployment probes. Limitations: model answers can be mistaken; no calendar booking or arbitrary code execution.
A loyalty portfolio with trip goals, qualified redemption estimates and reviewable agent proposals across web, MCP and browser-extension workflows.
PointUp organizes loyalty memberships, balances, expiry context and trip goals. Its public source defines a framework-independent TypeScript domain core shared by the Next.js web app, workers, MCP, typed API clients and browser extension. Supported balance capture requires consent; transfer and redemption values are qualified estimates, not bookable prices. The modern source does not by itself establish completion of the legacy application’s data migration.
A reproducible synthetic commerce lab for event validation, SQL metrics, dependency DAGs, graphs and vector similarity.
Built a reproducible synthetic commerce pipeline with seeded event generation, deduplication, schema and lifecycle checks, and SQL analytics. Explore acquisition and retention scenarios, inspect quarantined records, and trace conversion, collected revenue, and cohort retention to the queries that compute them. The public React lab serves generated Python runs without a database; a bounded FastAPI service supports custom simulations when configured. A separate synthetic commerce dataset connects products, customers, and purchases in a relationship graph and explains cosine similarity over handcrafted product feature vectors. An engineering workbench exposes executable dependency DAGs, retry and failure traces, model grains and contracts, and architecture decisions grounded in the implementation. The original PostgreSQL and Streamlit experiment remains in the source repository.
A research atlas for public data-center infrastructure, with source provenance and transparent uncertainty.
Built an infrastructure exploration product around public, sourced data. The atlas emphasizes what each source supports, where coverage is incomplete, and how to distinguish evidence from estimates. The public product is the available evidence; this card does not claim comprehensive global coverage.
A live-music companion for discovering concerts, making plans with friends, and keeping a show diary.
A public live-music product connecting concert discovery with plans and show memories. Public-facing copy supports the discovery, group-planning and diary narrative. Listing coverage and signed-in workflows have separate acceptance requirements.
A job-search workspace bringing job discovery, saved opportunities, and application work together.
A public job-search workspace for job discovery and career workflows. The Jobdog entry redirects to its Jobbr application. This product tile and the Jobbr entry describe the same public product family, not two independently verified deployments.
The Jobbr entry into Jobdog’s job-search and résumé workflow.
Jobbr is available inside Jobdog at the linked application. The older portfolio-hosted Jobbr demonstration uses synthetic browser-local data and is separate from this public product. Authenticated matching, preparation and application workflows require their own acceptance checks.
A prototype for per-thread AI token metering and cost reporting through an event-driven billing pipeline.
A React chat workspace connects to FastAPI over WebSocket with an HTTP send fallback. A separate collector, Kafka, Redis and SQLite support usage and billing events; Streamlit exposes cost analytics. The development workspace uses a seeded demo identity and has no authentication. Invoice creation exists in the backend but is unavailable in the React interface. The portfolio demo is a separate synthetic browser illustration; neither it nor the source establishes provider-backed billing acceptance.
A classroom copilot that answers 'who needs me today, and why?' with a ranked needs-attention list, a gradebook, attendance, and a streaming Claude assistant grounded in the class data.
Super Teacher is a classroom copilot for teachers. A Today view ranks the students who need attention with plain-language reasons; a roster, gradebook, and one-tap attendance feed the same computed averages, trends, and risk, so every screen agrees. Reports cover per-assignment class statistics, a CSV gradebook export, and AI-drafted parent updates the teacher edits before sending. An Ask AI chat streams from Claude with tools that query the real roster data, and per-student insight cards fall back to rule-based text when no API key is set. Built with a FastAPI backend on SQLAlchemy 2 and SQLite with Alembic migrations, and a React and Vite front end.
QR-code invitations for Facebook groups, covered in TechCrunch’s announcement of group-admin tools.
Contributed QR-code group invitations while at Meta. The linked public article documents the feature announcement; it does not quantify this contribution’s adoption or independently establish individual ownership.
Legacy Python and shell helpers for exporting Go-project context and running external lint, test and assistant workflows.
goPilot exports Go source context and invokes external Go checks through Python and shell helpers. Historical provider-aware commands use legacy assistant operations; current provider compatibility is not established. The project is not a Go application and has no application frontend. Its context-export and HTML-fetch helpers have offline tests; the portfolio demo is a separate synthetic illustration.
A historical Python project exploring market-data ingestion and query/regression experiments.
A 2017–2018 learning project with an alpha ingestion runner and omega query/regression experiments. The public source does not establish trade execution, a frontend or a production deployment. It is unmaintained and not runnable as-is. The portfolio’s synthetic browser demo is an explanatory illustration, not a live trading system.
A design-stage graph database for agents, connecting knowledge, tasks, actions, and provenance in one typed model.
Explores a shared, event-sourced graph for agent applications, with typed entities, attributable changes, and context retrieval. Current work includes architecture research and early storage foundations.
A Rust coding-agent engine and embeddable SDK built around a small core and replaceable providers, tools, and hosts.
Developing a shared execution engine for command-line and embedded coding agents. The design emphasizes explicit tool permissions, durable sessions, and extension points for providers and MCP tools.
A local control plane for coordinating coding agents, shared context, tools, and task handoffs across multiple harnesses.
Brings agent management, task context, tool access, and execution evidence into a local workspace. A working continuity service supports cross-harness handoffs while the broader management platform is in active development.
Local task-readiness and review tools that help coding agents clarify requirements and check results against an agreed contract.
Helps turn a request into a cited brief, explicit review, and machine-checkable acceptance criteria. A local task desk and coding-agent adapters keep the agreed scope and verification evidence together.