Hi, I'm Anna, a Principal Data Engineer based in New York City. I work at a growth marketing company, building the data platforms behind enterprise clients that spend billions on marketing.
My work is below. If you'd like to go deeper, architecture shows how I design these systems, and the terminal lets you query this site in SQL, Python, Java, or C.
also studieddata structures and algorithms · computer systems, C, and assembly · discrete math, proofs, and probability · object-oriented design in Python and Java
- Owned production data infrastructure supporting $2B+ in annual marketing spend, translating business requirements into scalable data pipelines, warehouse models, and reporting solutions.
- Built and deployed self-service automation tools and LLM-powered internal apps to reduce manual workflows, including a Raw Data Request App (AWS Amplify, Python Lambda, Docker, Snowflake, S3) that enables teams to configure and export datasets independently.
- Developed automated Slack alerting bots to proactively flag KPI anomalies, data changes, and taxonomy issues, helping teams identify and resolve problems before they impact reporting.
- Architected end-to-end data pipelines and API integrations across 50+ platforms for enterprise clients, including OpenAI, standardizing ingestion into Snowflake and BigQuery, transformation, QA, and reporting.
- Introduced AI-assisted engineering workflows using Claude Code and Codex, incorporating reusable skills, plugins, and agent harnesses to streamline SQL/Python development, QA, and technical documentation.
- Led cross-functional technical projects across engineering, client strategy, and activation teams, driving requirements, delivery planning, documentation, and operational standards.
- Mentored junior engineers in data modeling, pipeline debugging, documentation, and engineering best practices.
- Managed high-performance Tableau dashboards and automated Google Sheets reporting workflows, owning the full data lifecycle from ingestion and transformation through visualization and stakeholder delivery.
- Led client-facing data projects in partnership with Client Services, translating business requirements into technical scope, source-system definitions, cleaned datasets, and dashboards/reports that supported strategic decisions.
- Architected and implemented automated Google Sheets workflows for Client Services, enabling one-click input submission to trigger Apps Script and Matillion jobs that generated SQL views and wrote outputs back to designated tabs, reducing manual work and turnaround time for ad hoc data pulls by 80%.
- Refactored 50%+ of Snowflake data views to improve query performance, maintainability, and downstream reporting reliability.
- Monitored production alerts, investigated failures, and resolved data-quality issues to maintain accurate and reliable BI deliverables.
- Developed Snowflake SQL views powering automated reports and Tableau dashboards for 100+ clients.
- Designed ETL/ELT pipeline infrastructure integrating APIs, SFTP, and cloud storage into Snowflake via Matillion.
- Built Python ingestion scripts for custom JSON endpoints, using pandas for transformation and data cleanup.
- Built automated Slack alerting pipelines with SQL, Python, and Matillion to transform and aggregate complex datasets into client-ready insights, reducing issue resolution time by 40%.
- Resolved 200+ Jira tickets involving data discrepancies, failed integrations, and downstream reporting issues for Client Services teams.
- Reduced Snowflake warehouse costs by 30% by auditing 1,000+ tables/views, removing unused assets, and implementing view-to-table optimization workflows for high-cost reporting models.
- Wrote Confluence documentation to improve operational clarity and cross-team collaboration.
- Ingested data from multiple platforms and APIs into Domo, maintaining datasets, ETL dataflows, and dashboards that translated campaign performance data into actionable insights.
- Built 50+ interactive visualizations in Domo and automated the insights-generation process in Python using NumPy, pandas, and Matplotlib, reducing manual work and supporting monthly reporting and new-business pitches.
- Built Brandwatch queries and dashboards and created 10+ social listening and audience insights decks to support new business pitches.
- Wrote SQL to configure BigQuery API connectors and custom calculated fields in Domo for reporting.
- Performed QA on floodlight tags, tracking URLs, and UTM parameters, and resolved ad hoc data collection, taxonomy, and match-table issues to improve reporting accuracy.
The problem: nine ad platforms and a client's first-party data describe the same campaigns with different schemas, IDs, and names, and all of them change over time. This model, a simplified version of one I own, merges them into one daily table that's correct as of every date, without losing or double counting anything. Click any table to trace its lineage and SQL.