Data engineering is one of the fastest-growing and highest-compensating technology disciplines in the US right now. The demand for skilled data engineers significantly exceeds the supply — particularly for engineers who understand both the engineering and the data architecture sides of the role, not just the tools.
The growth is being driven by three forces simultaneously. First, consumer internet and fintech companies — DoorDash, Stripe, Airbnb, Brex, Robinhood, Instacart — have scaled to tens of millions of users and are now generating data volumes that require serious engineering to handle. Second, large enterprises — JPMorgan, Walmart, Amazon, Microsoft, Google — are building massive internal data platform teams, often paying well above the market median for senior talent. Third, the AI and ML wave has increased demand for the data pipelines that feed ML models — every company building AI features needs data engineers to prepare the training and inference data.
35,000+
DE job openings in the US
Active listings, March 2026
2.6×
Demand vs supply ratio
Skilled DEs vs open roles
14%
YoY salary growth
Mid-level DE, national median
$130K–$175K
Mid-level DE range
Product company, national median
4–8 months
Time to first job
From non-CS with the right prep
61%
Roles prefer cloud cert
AWS, Azure, or GCP cert
💡 Note
Data source: Salary figures in this module are sourced from Levels.fyi, Glassdoor, LinkedIn Salary Insights, and BLS Occupational Employment data, cross-referenced with data engineering community surveys. All figures reflect March 2026 data. Base salaries only — bonus and equity (RSUs) add 15–45% at product companies and public tech companies, and can add significantly more at pre-IPO startups if the exit is favorable.
// Part 02 — Real Salary Data
Salaries — What Data Engineers Actually Earn in the US
Salary data for data engineering in the US is scattered and often misleading — job boards conflate data analyst, data scientist, and data engineer salaries, and the ranges are wide enough to be unhelpful without context. Here is the breakdown by experience level, city, and company type with enough specificity to be genuinely useful for career planning.
By experience level — national median, product company baseline
DE salary by experience — product company, national median (2026)
Level Years Base Salary Range Total Comp (with equity)
──────────────────────────────────────────────────────────────────────
Junior DE 0–2 yrs $75K–$100K $80K–$110K
Entry into DE from
non-CS or CS new grad
Data Engineer 2–4 yrs $100K–$140K $110K–$155K
Owns pipelines end-
to-end independently
Senior DE 4–7 yrs $140K–$185K $155K–$215K
Designs systems,
mentors, cross-team
Staff / Lead DE 7–10 yrs $185K–$240K $215K–$290K
Technical strategy,
platform decisions
Principal DE 10+ yrs $240K–$320K+ $290K–$420K+
Company-level data
platform vision
Notes:
→ These are base salary ranges at well-paying product companies
→ Consulting firms (Accenture/Deloitte) pay 25–35% below these ranges
→ Large enterprises (JPMorgan, Walmart, banks) pay 10–20% above
→ FAANG (Amazon, Google, Meta) pay 55–90% above via base + RSU value
→ Equity at funded startups can add $20K–$300K+ in value if the exit lands
City multipliers — how location affects salary
The San Francisco Bay Area pays the most for data engineering in the US and is the reference point for all comparisons. Other cities pay varying multiples of the Bay Area base depending on the density of tech companies and local cost of living.
SF Bay Area
1.30×
Highest density of product companies and FAANG HQs
$150K base → $195K
Seattle
1.20×
Amazon, Microsoft, and a growing GCP presence
$150K base → $180K
New York City
1.20×
Strong fintech sector (Stripe, Robinhood, Brex)
$150K base → $180K
Austin
1.05×
Fast-growing hub, Tesla/Apple/Google offices
$150K base → $157K
Boston
1.05×
Mix of biotech, fintech, and established tech
$150K base → $157K
Chicago
0.95×
Reference city for enterprise benchmarking
$150K base → $142K
Remote (US)
1.10×
Many companies pay a premium to hire nationally
$150K base → $165K
Lower COL Metros
0.85×
Columbus, Raleigh, Tampa — growing but limited
$150K base → $127K
Company type multipliers — the biggest salary driver
Company type has a bigger impact on salary than city. The difference between working at a consulting firm and a FAANG operation is often 2–3× for the same role, experience, and city.
Salary multiplier by company type — applied to national mid-level base
Company Type Multiplier Mid-level Example Why
──────────────────────────────────────────────────────────────────────
FAANG / AI Labs 1.75× $175K–$260K Stock + high base,
(Amazon, Google, competitive global
Meta, OpenAI) talent market
Large Enterprise 1.15× $150K–$185K Stable pay bands,
(JPMorgan, Walmart, large-scale platform
Banks, Insurance) work
High-Growth Startup 1.10× $115K–$150K Equity adds value,
(Brex, Instacart, high learning rate,
Stripe, Robinhood) higher risk
Product Company 1.00× $130K–$175K Benchmark —
(Mid-size, funded) DoorDash, Shopify,
Airbnb, Databricks
Enterprise Software 0.95× $125K–$165K IBM, Oracle, SAP,
(non-FAANG) stable but slower-
moving stacks
Consulting 0.70× $95K–$130K Accenture, Deloitte, KPMG,
(IT services) Cognizant — volume
hiring, lower pay
Note on consulting firms: While salary is lower, consulting firms provide
structured training, large enterprise client exposure, and a recognizable
brand on a resume. Many engineers start here and move to product companies
after 2–3 years.
// Part 03 — Who Is Hiring
Top Companies Hiring Data Engineers in the US (2026)
These are the companies with consistent, high-volume data engineering hiring in the US right now. They are grouped by category with notes on what the work actually looks like at each type.
High-Growth Product Companies
The highest-learning environments for data engineering. Fast-growing data volumes, modern stacks, real production problems. Equity can be valuable at pre-IPO companies.
DoorDash
DE, Analytics Eng, Data Platform
Spark, Kafka, Airflow, dbt, Snowflake
Strong data platform team, good mentorship
Airbnb
DE, Data Platform Eng
Spark, Airflow, Druid, Kafka
Mature data platform, high engineering bar
Shopify
DE, Analytics Eng
Spark, dbt, Redshift, Airflow
Fast-growing, significant data engineering investment
Stripe
DE, Data Infra
Spark, Kafka, ClickHouse, Airflow
Payments data at scale, real-time requirements
Brex
DE, Analytics Eng
dbt, Snowflake, Airflow, Kafka
Modern stack, strong engineering culture
Robinhood
DE, Data Eng
Python, PostgreSQL, Redshift, Kafka
Fintech, growing data teams
Instacart
DE, Data Platform
Kafka, Spark, BigQuery
Real-time supply chain and marketplace data
Databricks / Snowflake
DE, Analytics Eng
Their own platforms, Python, SQL
Building the tools other data engineers use
Large Enterprises
High absolute salaries at senior levels, with the most stable pay bands. Work on internal data platforms with access to enterprise-scale problems and often deep legacy systems to modernize.
JPMorgan Chase
DE, Data Platform, Quant Data Eng
Spark, Python, internal platforms
Finance data at global scale, compliance-heavy
Goldman Sachs
DE, Data Engineer
Slang (internal), Python, BigQuery
Proprietary tech stack, top-tier comp
Walmart Global Tech
DE, Data Platform Eng
Spark, Kafka, Hive, Azure
Retail data at massive scale, legacy + modern mix
Amazon (AWS/Retail)
DE, SDE-Data
AWS native, Redshift, Glue, Kinesis
AWS-first stack, data engineering at Amazon scale
Microsoft
DE, Data Eng (Azure)
Azure-native, Databricks, Synapse
Azure stack depth, Azure certification valued
Google
DE, Data Eng
GCP-native, BigQuery, Dataflow, Pub/Sub
GCP depth, SWE-like hiring bar
Consulting Firms
Lower salary but large data engineering teams with consistent hiring. Good for getting a first job and building structured experience before moving to product companies.
Accenture
Data Engineer, ETL Developer
Informatica, SQL, basic Azure/AWS
Volume hiring, structured training programs
Deloitte
Data Engineer, Big Data Eng
Hadoop, Spark, SQL, cloud basics
Structured training track, client placement
KPMG
Data Engineer, Analytics Dev
Azure/AWS, SQL, Talend
Large data practice, many enterprise clients
Cognizant
Data Engineer, BI Developer
SQL, SSIS, Azure, Power BI
Banking and healthcare client focus
PwC
Data Engineer, Cloud Data Eng
Azure Databricks, ADF, Snowflake
Client-facing cloud data engineering work
Analytics Consultancies and Niche Players
Work across multiple client industries. Faster exposure to different data problems. Often a stepping stone to product companies.
Palantir
Forward Deployed Eng / Data Eng
Python, SQL, custom platforms
Embedded consulting, proprietary training
ThoughtWorks
Data Engineer
Modern cloud, dbt, Airflow, Spark
Strong engineering culture, client delivery focus
Slalom / Sigmoid
Data Engineer
AWS/Azure, Spark, dbt
Mid-size, specialised data engineering practices
// Part 04 — Skills in Demand
What US Companies Actually Hire For — The Real Skill Map
Job descriptions list dozens of tools, but hiring managers actually filter on a much smaller set of core competencies. Here's what genuinely moves the needle in a US data engineering interview, ranked by how often it appears as a hard requirement rather than a "nice to have."
Tier 1 — Non-negotiable (appears in 90%+ of JDs)
Core skills every US DE role expects
SQL (Advanced) Window functions, CTEs, query optimization,
reading EXPLAIN plans. This is tested in almost
every interview loop, often live.
Python Clean, production-grade code — not notebooks.
Comfortable with OOP, testing, and packaging.
A Cloud Platform AWS, Azure, or GCP — pick one and go deep rather
(pick one) than shallow across all three. AWS has the
largest market share of DE job postings.
Data Modeling Star schema, dimensional modeling, understanding
when to normalize vs. denormalize for analytics.
Orchestration Airflow is still the market standard, though
Dagster and Prefect are gaining share at
modern startups.
Tier 2 — Strong differentiators (appears in 40–70% of JDs)
Skills that separate mid from senior candidates
Spark / Distributed Computing Processing data at a scale where a
single machine can't keep up.
dbt The default transformation layer at
modern data companies. Fast to learn,
expected at analytics-forward teams.
Streaming (Kafka/Kinesis/Flink) Increasingly expected for senior
roles as more companies move toward
real-time analytics.
Data Warehouse Internals Snowflake, BigQuery, or Redshift —
understanding partitioning, clustering,
and cost optimization, not just SQL.
CI/CD & Infrastructure as Code Terraform, Docker, GitHub Actions.
Senior DE roles expect you to own
deployment, not just write pipelines.
🎯 Pro Tip
Depth beats breadth. A candidate who can talk in real detail about one production Spark pipeline they built and debugged will beat a candidate who lists ten tools they've only used in tutorials. Interviewers probe for depth within the first two follow-up questions.
⌨️
Try this yourself
Take your own resume or portfolio and pick the single project you know best. Write out, from memory, five follow-up questions an interviewer could reasonably ask about it — then check whether you can actually answer all five in real depth. Any gap is worth closing before you list that project as a strength.
// Part 05 — Decoding Job Postings
How to Read a US DE Job Posting — What They Really Mean
US job postings use a specific vocabulary that isn't always literal. Knowing how to translate it saves you from over-filtering yourself out of roles you're actually qualified for, and from under-preparing for ones you're not.
"5+ years of experience required"
Often negotiable for strong candidates with 3+ years and a solid portfolio — especially at startups. Rigid at large enterprises and government-adjacent companies.
"Must be authorized to work in the US"
The company will not sponsor a visa for this role. This is a hard filter — read it literally. Roles that do sponsor usually say "sponsorship available" or list it explicitly in the benefits section.
"Bachelor's degree in CS or related field, or equivalent experience"
The "or equivalent experience" clause is real and increasingly common. A strong portfolio and demonstrated skill can substitute for the degree requirement at most product companies and startups — less so at large enterprises and government contractors.
"Fast-paced environment"
Expect ambiguity, shifting priorities, and less process than a large enterprise. Common at startups — not necessarily bad, but know what you're signing up for.
"Ownership mentality"
You will be expected to make decisions without a manager specifying every detail. Common at product companies and startups; less emphasized at large enterprises where process is more defined.
"Competitive salary + equity"
Always ask for the actual number and the equity details (RSUs vs. options, vesting schedule, current 409A valuation if a private company) before accepting an offer.
⚠️ Important
If you're an international student or on a visa that requires sponsorship, filter job postings for this explicitly before applying — it saves significant time. Many company career pages let you filter by sponsorship availability, and it's a fair, direct question to ask a recruiter in a first screening call.
// Part 06 — The Non-CS Path
Breaking Into Data Engineering From a Non-CS Background
A large share of working data engineers in the US did not start with a computer science degree. The path is well-worn — community college transfers, coding bootcamp graduates, and fully self-taught engineers all land data engineering roles every year. What matters is what you can demonstrably build, not what your degree says.
The three realistic entry paths
Non-CS entry paths, ranked by typical time-to-first-job
Path Typical Timeline Cost Notes
──────────────────────────────────────────────────────────────────
Self-taught + 6–10 months $0–$500 Slowest but
Portfolio Projects (courses) cheapest. Requires
strong self-discipline
and a genuinely
impressive portfolio.
Coding Bootcamp 4–6 months $10K–$20K Structured, includes
(data-eng or SWE- + job hunt job-search support.
focused) Look for placement
rate transparency.
Community College / 1–2 years $3K–$10K Slower but builds
Associate's Degree genuine CS fundamentals
and can transfer credit
toward a 4-year degree.
What actually gets you hired without a CS degree
Hiring managers at product companies and startups care about three things, in this order: can you demonstrably build a working data pipeline, can you explain the tradeoffs in your design decisions, and can you communicate clearly with non-technical stakeholders. A degree is a proxy for these things — it is not the only proxy.
A real, end-to-end project
Not a tutorial clone. Ingest real data (a public API or dataset), transform it, load it into a warehouse, and document your design decisions in a README.
GitHub with clean commit history
Recruiters and hiring managers do look. Clean code, meaningful commit messages, and a well-written README matter more than raw project count.
One deep project beats five shallow ones
A single pipeline you can discuss in real depth for 20 minutes in an interview is worth more than five weekend projects you can only describe at a surface level.
A relevant cloud certification
AWS or Azure fundamentals certifications signal genuine effort and give recruiters a concrete, verifiable data point when your resume otherwise lacks a CS pedigree.
🎯 Key Takeaways
✓A CS degree is one signal among several — a strong portfolio and cloud certification can substitute at most product companies and startups.
✓Community college and bootcamp paths are well-worn and respected if the work produced is genuinely strong.
✓One deep, well-documented project beats five shallow tutorial clones in every interview.
✓Large enterprises and government-adjacent companies are more rigid about degree requirements than startups and mid-size product companies.
// Part 07 — Certifications
Which Certifications Actually Matter in the US (2026)
Certifications are not a substitute for real project experience, but they are a fast, verifiable signal — especially valuable if you're coming from a non-CS background or switching careers and don't yet have a work history in tech to point to.
Certifications ranked by hiring-manager relevance for DE roles
Certification Relevance Notes
──────────────────────────────────────────────────────────────
AWS Certified Data Engineer High AWS has the largest market
– Associate share of DE job postings.
Directly relevant content.
Microsoft Azure Data Engineer High Valuable specifically at
Associate (DP-203) companies already on Azure
(common in large enterprises).
Google Cloud Professional Medium-High Smaller market share than
Data Engineer AWS/Azure but growing,
especially at AI-forward
companies.
Databricks Certified Medium Valuable if the target
Data Engineer Associate company uses Databricks —
increasingly common.
dbt Fundamentals Low-Medium Free, fast to complete,
signals familiarity with
the modern data stack.
🎯 Pro Tip
Pick the certification that matches the cloud platform used by the companies you're actually targeting — check their job postings and engineering blog before choosing. A cert for a platform you'll never use professionally wastes study time better spent on a portfolio project.
⌨️
Try this yourself
Pull up the job postings of three companies you'd actually want to work for and tally which cloud platform (AWS, Azure, GCP) appears most often. Compare that against whichever certification you were already planning to pursue — if they don't match, that's worth reconsidering before you spend weeks studying for the wrong one.
// Part 08 — Negotiation
Salary Negotiation for Data Engineers in the US — The Honest Guide
US tech salary negotiation is different from most of the world in one key way: it's expected, and companies build room into their initial offer anticipating a counter. Not negotiating is the single most common way candidates leave money on the table.
The core rules
Negotiation principles that actually work
1. Never give the first number.
If asked for salary expectations, redirect: "I'd love to learn
more about the role first — what's the budgeted range?"
2. Get competing offers before negotiating.
A single offer gives you almost no leverage. Two offers, even
if one is less exciting, gives you real negotiating power.
3. Negotiate total compensation, not just base.
Base salary, signing bonus, equity (RSU count and vesting
schedule), and annual bonus target are all negotiable
independently. A lower base with more equity can be worth
more — or less — depending on the company's trajectory.
4. Use data, not feelings.
Reference Levels.fyi and Glassdoor figures for the specific
company and level. "Based on public data for this level at
your company, I was expecting closer to $X" is far stronger
than "I think I deserve more."
5. Get it in writing before you resign your current job.
Verbal offers can and do change. Wait for the signed offer
letter.
📌 Real World Example
A candidate with a competing offer at $145K received an initial offer of $155K base from their target company. By citing the competing offer and Levels.fyi data for the role and level, they negotiated the final offer to $172K base plus an increased equity grant — a 24% increase over the intial offer, achieved in a single negotiation email.
// Misconceptions
Five Misconceptions About the US Data Engineering Job Market
✕ ""You need a CS degree to get hired as a data engineer in the US""
Part 06 is explicit that a large share of working data engineers did not start with a CS degree — the "or equivalent experience" clause in job postings (Part 05) is real and increasingly common at product companies and startups. A strong portfolio and demonstrated skill substitute for the degree at most non-enterprise employers.
✕ ""The salary number in a job posting or on Glassdoor is basically fixed — there's not much room to negotiate""
Part 08 is direct that US tech salary negotiation is expected, and companies build room into their initial offer anticipating a counter — the worked example shows a 24% increase achieved in a single negotiation email. Not negotiating is called out as the single most common way candidates leave money on the table.
✕ ""City determines your salary more than anything else — just move to San Francisco or Seattle""
Part 02's "Company type multipliers" section is explicit that company type has a bigger impact on salary than city — the FAANG-vs-consulting gap for the same role, experience, and city is often 2-3×, larger than any city multiplier in that same Part.
✕ ""Listing as many tools as possible on your resume maximizes your chances""
Part 04's Callout and this module's first TryThis both make the opposite case: interviewers probe for depth within the first two follow-up questions, and a thin list of tools you can discuss deeply beats a long list that collapses under questioning.
✕ ""If you're getting rejected or hearing silence, it means you're not qualified""
Part 09's real career story reports roughly 100+ applications and single-digit interview conversion as the realistic range for a disciplined non-CS candidate — not evidence of being unqualified. This module's Common Mistakes and Error Library sections both cover the structural reasons (resume tailoring, ATS filtering) that explain most of this gap.
// Part 09 — Real World
From Retail Manager to Data Engineer — A Real Career Story
Composite story based on common patterns among Chaduvuko learners who broke into data engineering from a non-technical background.
Starting point: A retail store manager with a bachelor's degree in Business Administration, no programming background, working 60-hour weeks and looking for a career with better work-life balance and higher earning ceiling.
Months 1–3: Learned Python and SQL fundamentals through free resources during evenings and weekends. Built the first small project — a script that pulled data from a public API and loaded it into a local SQLite database. Unimpressive by portfolio standards, but it proved the core loop was learnable.
Months 4–7: Enrolled in a part-time, remote data engineering bootcamp while still working retail. Built two more substantial projects: an Airflow-orchestrated pipeline pulling weather data into a cloud data warehouse, and a dbt project transforming raw e-commerce data into analytics-ready tables. Got AWS Cloud Practitioner and then AWS Data Engineer Associate certified.
Months 8–9: Applied to roughly 120 positions. Most were silence or rejection. Landed 6 first-round interviews, largely from roles where the JD explicitly said "or equivalent experience." Two progressed to technical rounds. One offer came through — Data Engineer at a mid-size logistics startup, $92K base, fully remote.
Where they are now (18 months in): Promoted once, now earning $118K base plus equity. Actively interviewing for senior roles at larger product companies using the same portfolio-first approach that got the first job.
💡 Note
The honest numbers: roughly 100+ applications, single-digit interview conversion, and 8–9 months from starting to learn to first offer. This is the realistic range for a disciplined non-CS candidate, not the outlier "landed a job in 6 weeks" stories that circulate online.
// Part 10 — Interview Prep
5 Interview Questions — With Complete Answers
1
How would you design a pipeline to ingest 10 million events per day from a third-party API into a data warehouse?
Start by clarifying requirements: latency needs (batch vs. near-real-time), data volume growth expectations, and downstream consumers. For a batch approach: use an orchestrator (Airflow) to poll the API on a schedule, land raw data in object storage (S3) as the source of truth, then load into the warehouse via a managed loader. For near-real-time: consider a message queue (Kafka/Kinesis) if the API supports webhooks or streaming, with a stream processor writing to the warehouse in micro-batches. Always discuss idempotency (handling retries without duplicating data), schema evolution (the API will change), and monitoring (alerting on ingestion lag or failure).
2
Explain the difference between a data lake and a data warehouse, and when you'd use each.
A data warehouse stores structured, schema-on-write data optimized for fast analytical queries — used when you know your query patterns in advance and need strong performance and governance (BI dashboards, reporting). A data lake stores raw, often unstructured or semi-structured data with schema-on-read — used when you need flexibility to store diverse data types cheaply and figure out the schema later (ML training data, exploratory analysis, archival). Most modern architectures use both together (a "lakehouse" pattern, e.g. Databricks or a lake feeding a warehouse) — raw data lands in the lake, gets cleaned and modeled, then loads into the warehouse for consumption.
3
How do you handle schema changes in a production pipeline without breaking downstream consumers?
Prefer additive, backward-compatible changes (new nullable columns) over breaking changes (renaming or removing columns, changing types). Use a schema registry if working with streaming data (Kafka + Avro/Protobuf) to enforce compatibility rules automatically. For batch pipelines, version your schemas and use tools like dbt's contracts or Great Expectations to validate incoming data against an expected schema before it propagates downstream. Communicate breaking changes to consumers in advance with a deprecation window, and maintain both old and new schema versions during the transition if the consumer base is large.
4
A daily batch job that used to take 20 minutes now takes 3 hours. How do you debug it?
Start with what changed: data volume growth, a code change, a cluster configuration change, or an upstream data quality issue causing unexpected joins/skew. Check the query plan (EXPLAIN) for the slowest stage of the pipeline — look for data skew (one partition doing disproportionate work), unnecessary shuffles, or a missing partition filter causing a full table scan. Check cluster metrics (CPU, memory, spill-to-disk) to rule out resource contention. If using Spark, check the Spark UI for stage-level timing. Fix the root cause rather than just scaling up the cluster, which masks the problem and increases cost.
5
How do you ensure data quality in a pipeline that multiple teams depend on?
Implement automated checks at ingestion (schema validation, null checks, range checks) and post-transformation (row count reconciliation, referential integrity, freshness checks) using a framework like Great Expectations or dbt tests. Set up alerting so failures page someone before bad data reaches dashboards, not after. Maintain a data catalog with clear ownership so consumers know who to contact. For critical pipelines, consider a staging layer where data is validated before being promoted to production tables, so a bad run never silently corrupts the tables everyone queries.
// Common Mistakes
Mistakes You Will Make — And Exactly Why They Happen
Applying to 200+ jobs with the same generic resume
Applicant tracking systems and human reviewers both filter fast. A resume tailored to the specific JD's keywords and requirements converts at a meaningfully higher rate than a one-size-fits-all version.
Listing every tool you've ever touched instead of what you can actually discuss in depth
Interviewers ask follow-up questions on anything listed. A thin list of tools you can discuss deeply beats a long list that collapses under two follow-up questions.
Not clarifying work authorization status upfront if it requires sponsorship
This wastes both your time and the recruiter's if discovered late in the process. Address it directly and early — most experienced recruiters appreciate the directness.
Negotiating with emotion instead of data
"I really need this" is a weak negotiating position. "Based on public data for this level, the market rate is X" is a strong one. Always negotiate with a specific number and a specific source.
Underestimating how long the US job search actually takes
Median time from active search to signed offer for a career switcher is realistically 3–6 months, not the 4–6 week timelines that circulate in viral success stories. Budget your finances and expectations accordingly.
Skipping the behavioral interview prep because "it's just the technical round that matters"
Most US tech companies weight behavioral rounds heavily, often as a hard gate regardless of technical performance. Prepare specific, structured stories (the STAR format) in advance — don't improvise them live.
// Error Library
Job Search Breakdowns — And Exactly Why They Happen
Application submitted through the company careers page — no response, no rejection email, complete silence after three weeks
Cause: The resume was filtered out by the Applicant Tracking System (ATS) before any human reviewed it. Most large-company ATS software scores resumes against the exact keywords in the job description — a resume that says "data pipelines" when the JD says "ETL pipelines" can score low enough to never surface, even from a genuinely qualified candidate.
Fix: Mirror the exact terminology from the job description in the resume (without lying about experience) — if the JD says "data warehousing," use that phrase, not a close synonym. For roles at companies known to use aggressive ATS filtering, apply through a referral or direct recruiter message in parallel with the formal application, since a human-forwarded application usually bypasses the ATS scoring step entirely.
Recruiter phone screen ends abruptly right after the candidate states their salary expectation
Cause: The stated number was either far below the role's actual level (signaling the candidate may be under-qualified or a poor culture fit for a more senior title) or far above the budgeted range for that specific req — either way, the recruiter has no room to continue the conversation productively.
Fix: Per Part 08, never give the first number — redirect and ask for the budgeted range instead. If pressed, research the specific level and company on Levels.fyi beforehand and give a wide, well-researched range rather than a single guessed figure, so the number reflects the actual role rather than a generic industry average.
Technical interview goes well, but the process goes silent after the take-home assignment is submitted
Cause: Most take-home assignments are evaluated primarily on code quality, testing, and documentation — not just whether the output is correct. A working solution with no tests, no error handling, and no README explaining design decisions reads as unfinished, even if the core logic is right.
Fix: Treat every take-home as a portfolio piece, not just a puzzle to solve: include basic tests, handle at least the obvious edge cases, and write a short README explaining the design trade-offs made under the time constraint. This directly reflects the "explain the tradeoffs in your design decisions" criterion from Part 06.
A verbal offer is extended over a call, but the candidate resigns from their current job before receiving anything in writing — then the signed offer never arrives
Cause: Verbal offers can and do change or fall through — a budget freeze, a hiring-manager change, or an internal approval that never finalizes can all cause a verbal offer to quietly disappear, and without a resignation already submitted, the downside is limited.
Fix: Part 08's negotiation rules are explicit: get the offer in writing before resigning a current job. This is not paranoia — it is standard practice, and no reputable company will consider it an unreasonable request from a candidate.
A candidate on a work visa gets to the final round, then the process ends with "we've decided to go a different direction" with no further explanation
Cause: Sponsorship requirements were not clarified early, and the company either does not sponsor visas for this role or has hit an internal cap on sponsorship approvals for the year — a constraint that has nothing to do with the candidate's technical performance, but surfaces only at the final stage when it becomes relevant to the hiring paperwork.
Fix: Per Part 05's Callout, filter for sponsorship availability before applying wherever possible, and ask the recruiter directly and early in the first screening call. This does not need to feel adversarial — most experienced recruiters would rather resolve this upfront than run a candidate through four interview rounds only for it to become a blocker at the end.
What comes next
Module 07 covers the three data categories every data engineer works with daily — structured, semi-structured, and unstructured — and what each one demands from your pipeline design.