~/anna-vu
SELECT *  FROM  anna_vu.creative_projects  ORDER BY year  DESC -- personal projects · 10 rows returned
WHERE project_name LIKE '% %'
AND category =
#PROJECT CATEGORYSTATUS
01
Film Festival Tracker Track documentary and experimental film festivals worldwide. Deadlines, fees, and acceptance rates, filterable by type and region. view project →
Film
live
02
Analog Film Stock DatabaseCompare analog film stocks by price, availability, and format. Paired with a development lab directory searchable by process and location.view project →
Film
live
03
NYC Film Scene MapInteractive map of NYC screening spaces, workshops, labs, galleries, and community film spots. Filterable by category and borough.view project →
Film
live
04
AuteursPersonal index of filmmakers and performers, searchable by era, nationality, and style. Includes film guides and portraits.view project →
Film
live
05
NYC Food Help MapPublic-service directory of free food assistance organizations across NYC: pantries, soup kitchens, community meals, mobile food, and home delivery. Filterable, mapped, and volunteer-linked.view project →
Civic
live
06
NYC Adventure LogCurated NYC activities across seasons and neighborhoods. Cinema, Food, Arts, Explore, and more. Surprise picker included.view project →
City
live
07
Only Good ThingsCurated self-care activity tracker with filters for season, duration, mood, and setting. Tracks completions locally, with a daily pick to nudge you toward trying one.view project →
Wellness
live
08
Cơm NhàVietnamese recipes from Miền Bắc, Miền Trung, and Miền Nam. Browse by region, search by ingredient.view project →
Food
live
09
Departure LoungeDestinations as boarding passes, with a calendar and budget planner for actually booking the trip.view project →
Travel
live
10
Global Donation GuideCurated directory of humanitarian relief, global hunger, and animal welfare organizations worldwide. Each one researched individually, with direct donation links.view project →
Giving
live
SELECT project_name, company, year, stack  FROM  anna_vu.projects  ORDER BY year  DESC -- 14 rows · engineering + academic

Hi, I'm Anna, a Principal Data Engineer based in New York City. I work at a growth marketing company, building the data platforms behind enterprise clients that spend billions on marketing.

My work is below. If you'd like to go deeper, architecture shows how I design these systems, and the terminal lets you query this site in SQL, Python, Java, or C.

Data Pipelines & Client Dashboard Suite
2022–Present
SnowflakePythonMatillionLookerTableau
A major enterprise marketing client needed one trustworthy view of performance across every platform. I own the infrastructure end to end: 50+ sources ingested into Snowflake and modeled once, powering Tableau and Looker dashboards, automated reports, and ad-hoc pulls for 100+ clients.
50+
sources
100+
clients
$2B+
ad spend
Raw Data Request Builder
2026
AWS AmplifyLambdaDockerSnowflakePython
Client Services used to wait on engineering for every raw data pull. Now they pick the data and get the file themselves. Amplify frontend, a Dockerized Python Lambda running parameterized Snowflake queries, CSV/Excel/JSON exports staged to S3 and delivered to Google Drive, and an audit log written back to Snowflake on every request.
3
export formats
S3 + Drive
delivery
Zero-code
for end users
Automated Slack Alerting Suite
2022–Present
SQLPythonSnowflakeMatillionSlack API
Data issues used to surface when someone spotted a broken dashboard. These production alerting pipelines catch them first and route them to the right team in Slack: SFTP file arrivals, KPI anomalies across platforms and geos, changes in reporting docs, and taxonomy/match-table checks that flag bad ad naming to the Media team before it reaches reporting.
4+
alert types
Multi-team
routing
Config-driven
thresholds
data engineering
Snowflake Warehouse Optimization
2023–2024
SnowflakeSQL
Years of ungoverned growth left the warehouse full of duplicate views, inconsistent naming, and objects nobody owned. I audited 1,000+ tables and views, consolidated redundant logic, standardized naming, and removed unused objects. Compute costs dropped 30%.
30%
cost reduction
1,000+
tables audited
Google Sheets Data Pull Tool
2024
Apps ScriptMatillionSnowflakeGoogle SheetsSQL
Teams worked in Google Sheets but had to request Snowflake data by hand. I built self-serve reporting inside Sheets: choosing dimensions and metrics triggers Apps Script and Matillion pipelines that generate Snowflake views and write results back to the sheet. Turnaround time dropped 80%.
80%
less manual work
80%
faster turnaround
Custom API Ingestion & Data Processing
2022–2023
PythonPandasREST APIsSnowflake
Every new marketing platform used to mean a connector built from scratch. I built modular Python connectors for 6+ platforms (Google Ads, CM360, Meta, DV360) that handle OAuth, pagination, and rate limits and load into Snowflake, so onboarding a new source takes hours, not days.
6+
platforms
days → hrs
onboarding
earlier work
ESOV Media Spend Analysis
2022
RKantarPathmaticsGoogle Trends
Brand teams needed evidence for how much to spend in their annual budgets. I built the analysis in R on Kantar, Pathmatics, and Google Trends data: ESOV models, CDI/BDI scatter plots, and competitor salience benchmarks that fed annual budget strategy.
3
data sources
Competitive
intel framework
Social Listening & Audience Insights
2022
BrandwatchPythonJupyter
New business pitches needed audience insight, fast. I built Brandwatch social listening dashboards across client brands, plus Python scripts that pulled the data, generated charts, and drafted narratives for 10+ pitch decks.
10+
insights decks
Media Campaign Trafficking
2021–2022
Prisma/MediaOceanCampaign Manager 360Google Ads
Where I started: campaign operations across 20+ media vendors in digital, OOH, and print. IO setup, budget pacing, pixel trafficking, tag QA, and post-campaign analysis.
20+
vendors
Multi-channel
digital · OOH · print
penn engineering · projects
LC4 Toolchain: Assembler & Disassembler
2026
CISABinary I/OLinked ListsValgrindGDB
Both directions of the LC4 toolchain, in C. The assembler runs a two-pass parse over assembly source and encodes every instruction into 16-bit machine code in PennSim's .obj format, handling .CODE/.DATA/.SYMBOL sections and endianness. The disassembler reads those binary files back into a linked list that models program and data memory and prints readable assembly; written solo with manual memory management, it passes all 14 Valgrind memory checks.
30+
opcodes
14/14
valgrind checks
COVID & Property Analytics Platform
2025
JavaCSV/JSONMVCMemoizationJUnit
Data processing platform combining Philadelphia COVID vaccination records, population by ZIP, and property assessments. MVC architecture with custom readers, Singleton logger for audit trails, and memoized processors caching expensive computations.
3
data sources
7+
test suites
Flu Tweet Geolocation Analyzer
2025
JavaRegexGeolocationSingletonFactory
Analyzes geotagged tweets for flu-related content using regex pattern matching (handling edge cases like "fluent" vs "#flu"). Maps matches to the nearest U.S. state via Cartesian distance. Polymorphic Reader interface supports JSON and TXT formats.
50
states mapped
2
input formats
RFC 4180 CSV Parser
2025
JavaState MachineParsing
A CSV reader built as a four-state finite state machine that follows RFC 4180: quoted fields with embedded commas and line breaks, doubled quotes as escapes, CRLF and LF row endings, and malformed input rejected with an error that pinpoints the line and column.
4
parser states
RFC 4180
spec
Graph Traversal Algorithms
2025
JavaGraphsBFS
Analysis methods for directed and undirected graphs stored as adjacency sets: breadth-first search for the shortest distance between two nodes, every node within k hops, and a Hamiltonian cycle checker that reports exactly why a path fails (wrong length, not a cycle, or an unknown node).
3
algorithms
BFS
traversal

also studieddata structures and algorithms · computer systems, C, and assembly · discrete math, proofs, and probability · object-oriented design in Python and Java

SELECT role, company, period, wins  FROM  anna_vu.experience  ORDER BY start_date  DESC -- 4 records
[01]
Principal Data Engineer
DEPT® Agency · New York, NY · Aug 2025 – Present
AWS AmplifyLambdaDockerSnowflakeBigQueryPythonSQLMatillionS3Claude CodeCodex
  • Owned production data infrastructure supporting $2B+ in annual marketing spend, translating business requirements into scalable data pipelines, warehouse models, and reporting solutions.
  • Built and deployed self-service automation tools and LLM-powered internal apps to reduce manual workflows, including a Raw Data Request App (AWS Amplify, Python Lambda, Docker, Snowflake, S3) that enables teams to configure and export datasets independently.
  • Developed automated Slack alerting bots to proactively flag KPI anomalies, data changes, and taxonomy issues, helping teams identify and resolve problems before they impact reporting.
  • Architected end-to-end data pipelines and API integrations across 50+ platforms for enterprise clients, including OpenAI, standardizing ingestion into Snowflake and BigQuery, transformation, QA, and reporting.
  • Introduced AI-assisted engineering workflows using Claude Code and Codex, incorporating reusable skills, plugins, and agent harnesses to streamline SQL/Python development, QA, and technical documentation.
  • Led cross-functional technical projects across engineering, client strategy, and activation teams, driving requirements, delivery planning, documentation, and operational standards.
  • Mentored junior engineers in data modeling, pipeline debugging, documentation, and engineering best practices.
50+platforms
$2B+ad spend
OpenAIclient
3engineers mentored
[02]
Senior Business Intelligence Engineer
DEPT® Agency · New York, NY · Jun 2024 – Aug 2025
TableauGoogle SheetsApps ScriptMatillionSnowflakeSQLPython
  • Managed high-performance Tableau dashboards and automated Google Sheets reporting workflows, owning the full data lifecycle from ingestion and transformation through visualization and stakeholder delivery.
  • Led client-facing data projects in partnership with Client Services, translating business requirements into technical scope, source-system definitions, cleaned datasets, and dashboards/reports that supported strategic decisions.
  • Architected and implemented automated Google Sheets workflows for Client Services, enabling one-click input submission to trigger Apps Script and Matillion jobs that generated SQL views and wrote outputs back to designated tabs, reducing manual work and turnaround time for ad hoc data pulls by 80%.
  • Refactored 50%+ of Snowflake data views to improve query performance, maintainability, and downstream reporting reliability.
  • Monitored production alerts, investigated failures, and resolved data-quality issues to maintain accurate and reliable BI deliverables.
80%faster turnaround
50%+views refactored
100+clients
[03]
Business Intelligence Engineer
DEPT® Agency · New York, NY · Oct 2022 – Jun 2024
SQLMatillionPythonPandasSnowflakeTableauREST APIsSlack API
  • Developed Snowflake SQL views powering automated reports and Tableau dashboards for 100+ clients.
  • Designed ETL/ELT pipeline infrastructure integrating APIs, SFTP, and cloud storage into Snowflake via Matillion.
  • Built Python ingestion scripts for custom JSON endpoints, using pandas for transformation and data cleanup.
  • Built automated Slack alerting pipelines with SQL, Python, and Matillion to transform and aggregate complex datasets into client-ready insights, reducing issue resolution time by 40%.
  • Resolved 200+ Jira tickets involving data discrepancies, failed integrations, and downstream reporting issues for Client Services teams.
  • Reduced Snowflake warehouse costs by 30% by auditing 1,000+ tables/views, removing unused assets, and implementing view-to-table optimization workflows for high-cost reporting models.
  • Wrote Confluence documentation to improve operational clarity and cross-team collaboration.
100+clients
30%cost cut
200+tickets
40%faster resolution
[04]
Data Analyst
FIG Agency · New York, NY · Jun 2021 – Jun 2022
DomoPythonSQLBigQueryBrandwatchNumPyPandasMatplotlib
  • Ingested data from multiple platforms and APIs into Domo, maintaining datasets, ETL dataflows, and dashboards that translated campaign performance data into actionable insights.
  • Built 50+ interactive visualizations in Domo and automated the insights-generation process in Python using NumPy, pandas, and Matplotlib, reducing manual work and supporting monthly reporting and new-business pitches.
  • Built Brandwatch queries and dashboards and created 10+ social listening and audience insights decks to support new business pitches.
  • Wrote SQL to configure BigQuery API connectors and custom calculated fields in Domo for reporting.
  • Performed QA on floodlight tags, tracking URLs, and UTM parameters, and resolved ad hoc data collection, taxonomy, and match-table issues to improve reporting accuracy.
50+visualizations
10+pitch decks
5+API integrations
SELECT *  FROM  anna_vu.lineage -- anonymized model · generic names, no client data
Data platform · click any stage or component to see related projects
one reporting model, up close

The problem: nine ad platforms and a client's first-party data describe the same campaigns with different schemas, IDs, and names, and all of them change over time. This model, a simplified version of one I own, merges them into one daily table that's correct as of every date, without losing or double counting anything. Click any table to trace its lineage and SQL.

design decisions
CONNECT TO  anna_vu.profile  -- SQL · Python · Java · C · type a query below
anna_db> 
anna_vu :: snowflake UTF-8 studio 9 rows | 2.1ms
PUMKIN
z z Z