Data Engineer · Analytics Engineer

Donnoban Maldonado

AI-native engineer building data and analytics systems.
Based in NY.

Portrait of Donnoban Maldonado

Experience

My work.

View resume
  1. Customer Solutions Architect

    dbt Labs · New York, NY

    Apr 2026 — Jul 2026
    • Architected scalable dbt solutions for a $2.8M ARR portfolio of enterprise clients — CI/CD workflows, UAT environments, DAG dependency resolution, semantic layer adoption, and cost-optimized pipelines.
    • Drove enterprise platform adoption by analyzing product telemetry to proactively identify account risk and surface growth opportunities.
    • Automated account telemetry reviews with Claude, skill specs, and subagent orchestration — batching Salesforce, Omni, and Gong transcript data modeled in Snowflake through isolated subagents to synthesize account health, notes, and next steps for architect approval before syncing back to Salesforce. Saved architects 5+ hours weekly.
    • dbt
    • Snowflake
    • Semantic Layer
    • CI/CD
    • Claude Code
    • Salesforce
    • Omni
  2. Analytics Engineer

    Publicis Groupe · New York, NY

    Jul 2025 — Apr 2026
    • Built and maintained marketing reporting pipelines in dbt, BigQuery, and Git for the Verizon Value Brands portfolio (TracFone, Visible, Verizon Prepaid), feeding Tableau dashboards used by 500+ combined daily visitors.
    • Optimized pipelines processing billions of records with insert-by-period and incremental logic — ~6× faster daily runs and ~70% lower warehouse production costs.
    • Developed a net-new data product MVP to track and attribute previously unmeasured inbound campaign performance, delivering a portfolio-wide executive rollup.
    • Built regression testing on top of dbt snapshots to flag historical shifts in source data, catching upstream quality issues before they reached client-facing dashboards.
    • Audited the dbt test suite to eliminate noise and bridge coverage gaps, increasing stakeholder trust in the data products.
    • Owned an internal observability data product on the Elementary dbt package, automating alerting and visualizing pipeline runs and test failures in Looker.
    • Monitored daily health of production pipelines, triaging failures and deploying hotfixes before they disrupted stakeholder-facing dashboards.
    • Partnered with clients and cross-functional stakeholders to gather requirements and give technical guidance on pipeline capabilities and data.
    • dbt
    • BigQuery
    • SQL
    • Looker
    • Tableau
    • Elementary
    • Git
  3. Research Assistant

    SUNY Research Foundation · Old Westbury, NY

    Jan 2025 — Jul 2025
    • Designed a data collection system structuring Discord server data for machine learning in Python, Pandas, the Google Sheets API, and GCP — later deployed at Cal Poly and University of Texas Victoria.
    • Analyzed 23,000+ Reddit posts with a mixed-methods NLP pipeline, showing that LLMs outperform traditional models at extracting actionable insight on transfer student challenges.
    • Presented findings on pipeline design and NLP analysis at the SUNY Undergraduate Research Conference, hosted by SUNY Binghamton.
    • Python
    • Pandas
    • NLP
    • LLMs
    • GCP

Education

B.S. Computer Science

SUNY Old Westbury · Old Westbury, NY

Conferred May 2025
Honors
Summa Cum Laude
GPA
3.98

Relevant coursework

  • Machine Learning
  • Data Mining
  • Database Management Systems
  • Software Engineering

Capabilities

What I work with.

Languages & Tools

  • Python
  • SQL
  • dbt
  • Git
  • Pandas
  • Claude Code

Data Platforms

  • BigQuery
  • Snowflake
  • GCP
  • AWS RDS
  • MySQL

BI & Visualization

  • Looker
  • Omni
  • Tableau
  • Power BI
  • Google Sheets

Methodologies

  • Dimensional / Kimball
  • One Big Table
  • CI/CD
  • Test-Driven Development
  • Semantic Layer

AI & NLP

  • LLM Workflows
  • Subagent Orchestration
  • Topic Modeling
  • Word2Vec
  • Sentiment Analysis

Spoken

  • English
  • Spanish

Selected Work

Things I've built.

Terminal output from the job application logger

Sep 2026

Job Application Logger

A job search means keeping a spreadsheet, and maintaining one by hand is its own small job. This tool keeps the spreadsheet from the mail instead: one command pulls the job-application email that arrived since yesterday, classifies each message in-session with Claude, cross-checks it against the sheet, and prints a table. Nothing is written until you approve it. The tool is open source under an MIT license.

  • Node.js
  • Gmail API
  • Sheets API
  • Claude Code
Discord bot data collection project

May 2025

Discord Bot for Data Collection

A Python Discord bot built to support and study transfer students in computing. It builds peer networks while collecting anonymized engagement data to inform AI-driven advising. Hosted on GCP with the Google Sheets API for storage, it verifies users by university email, assigns roles, and logs activity behind hashed identifiers. Later deployed at Cal Poly and UT Victoria.

  • Python
  • GCP
  • Sheets API
  • discord.py
Magic Sheets worksheet generator

Jun 2025

Magic Sheets

A generative-AI web app that lets educators, parents, and students create customizable K–12 worksheets in seconds. Pick a subject, grade, topic, and format, then refine the output with prompts like "simplify" or "make it more descriptive." Built with Django and GPT-4o, with in-browser DOCX editing, answer key toggling, and a community ecosystem for sharing worksheets.

  • Django
  • Python
  • GPT-4o
  • PostgreSQL
Sentiment analysis project

May 2025

Film Review Sentiment Analysis

Classifies IMDb movie reviews as positive or negative using NLP and machine learning. After cleaning and vectorizing text with Word2Vec (Skip-gram), several classifiers were trained and compared. A two-layer Multi-Layer Perceptron over 500-dimensional embeddings performed best, reaching an F1-score of 0.8911.

  • Python
  • Word2Vec
  • scikit-learn
  • NLP
Tech layoffs interactive report

Oct 2024

Tech Layoffs Report

An end-to-end pipeline tracking tech industry layoffs. Raw CSV data was loaded into an AWS RDS MySQL instance and cleaned in SQL — deduplication, naming normalization, null handling — moving from a staging table to a final table before being visualized as a multi-page interactive Looker report covering trends, company breakdowns, and geographic impact.

  • SQL
  • AWS RDS
  • MySQL
  • Looker
Suffolk County salary dashboard

Aug 2024

Suffolk County Salaries

An interactive Tableau dashboard on public payroll spending in Suffolk County, NY. Data was cleaned and preprocessed in Python to strip sensitive fields and consolidate job titles. The analysis found that over 25% of total compensation went to the police department — more than every other department combined.

  • Python
  • Pandas
  • Tableau

Get in Touch

Tell me what you're building.

I'm currently open to data and analytics engineering roles. If you're building something interesting with data — or something that's stopped being interesting because the pipelines keep breaking — I'd like to hear about it.

  • Looking for Hands-on data & analytics engineering
  • Based in Long Island, NY — open to NYC and remote
  • Response time Usually within a day