DataHub
DataHub7h ago
New

Senior Software Engineer or Staff Engineer- Ingestion Framework

East CoastRemotesenior
Software EngineerSoftware Engineering
2 views0 saves0 applied

Quick Summary

Key Responsibilities

pt-0 [&>p]:mb-2 [&>p]:my-0"> Build and evolve our Python-based metadata-ingestion framework, enabling rich ingestion of usage statistics, end-to-end lineage, schemas, operational metadata,

Technical Tools
Software EngineerSoftware Engineering

DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility.

The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos.

Help build the ingestion engine behind DataHub’s trusted context layer for data and AI. You’ll make it easier for enterprises to connect the systems that power their business—capturing the metadata, lineage, usage signals, and operational context that people and AI agents need to make reliable decisions.

DataHub’s ingestion framework already powers metadata collection across a broad modern data stack; this role will expand and evolve that foundation with deeper integrations, cloud-native execution, and AI-assisted capabilities.datahub+1

Responsibilities

~1 min read
  • p]:pt-0 [&>p]:mb-2 [&>p]:my-0">

    Build and evolve our Python-based metadata-ingestion framework, enabling rich ingestion of usage statistics, end-to-end lineage, schemas, operational metadata, quality signals, and business context.

  • p]:pt-0 [&>p]:mb-2 [&>p]:my-0">

    Design and ship high-quality connectors for the modern data and AI ecosystem—including platforms such as Snowflake, Redshift, Kafka, Databricks, and emerging infrastructure.

  • p]:pt-0 [&>p]:mb-2 [&>p]:my-0">

    Make ingestion cloud-native: scalable, resilient, observable, and easy to deploy and operate across customer environments.

  • p]:pt-0 [&>p]:mb-2 [&>p]:my-0">

    Apply AI to create smarter connectors, including metadata inference, automated enrichment, and more intelligent ingestion workflows.

  • p]:pt-0 [&>p]:mb-2 [&>p]:my-0">

    Solve challenging distributed-systems problems involving large-scale metadata extraction, incremental updates, reliability, and performance.

  • p]:pt-0 [&>p]:mb-2 [&>p]:my-0">

    Lead technical direction for a small engineering pod, mentor junior engineers, and raise the bar for engineering quality, design, and execution.

  • p]:pt-0 [&>p]:mb-2 [&>p]:my-0">

    Work closely with product, platform, and customer-facing teams to turn real-world data-integration needs into polished product capabilities.

  • 8+ years of experience designing, building, and shipping production distributed backend systems.

  • Strong Python expertise, including hands-on experience owning features from design through implementation, testing, deployment, and iteration.

  • Excellent computer-science fundamentals, supported by a degree in Computer Science or equivalent practical experience.

  • Deep understanding of API design, algorithmic complexity, reliability, and distributed-systems concepts.

  • A practical, ownership-oriented mindset: you thrive in fast-moving environments, navigate ambiguity well, and turn difficult technical problems into durable solutions.

  • Experience with modern data platforms such as Snowflake, Databricks, Kafka, or similar systems is highly valued.

  • Experience operating across multiple cloud environments is a strong plus.

  • Previous experience leading or mentoring a small team or engineering pod is a strong plus.

This is an opportunity to build foundational infrastructure at the intersection of data, cloud, and AI. DataHub is an open-source metadata platform that brings together discovery, governance, observability, lineage, and real-time context across an organization’s data ecosystem—making that context actionable for both human teams and AI agents.datahub+1

You will have meaningful ownership over a product area that directly shapes how enterprises understand and trust their data:

  • Build products used across complex, modern enterprise data stacks.

  • Work on highly technical problems with direct customer and ecosystem impact.

  • Help define how AI agents gain safe, governed, and useful context about enterprise data.

  • Contribute to an open-source platform with a large global community and broad adoption.

What We Offer

~1 min read

We invest in people so they can do their best work and enjoy doing it. Our benefits reflect the way we build: practical, thoughtful, and designed to support long-term growth.

We offer salaries that reflect your skills, experience, and the impact you make. You bring value—we make sure you're recognized for it.

DataHub is at a rare inflection point: we’ve achieved product-market fit, earned the trust of leading enterprises, and secured backing from top-tier investors like Bessemer Venture Partners and 8VC. The context platform market is expected to grow from $1B to $9B in the next five years—and we’re leading the way.

By joining our team, you’ll:

Tackle high-impact challenges at the heart of enterprise AI infrastructure
Ship production systems that power real-world use cases at global scale
Collaborate with a high-caliber team of builders who’ve scaled some of the most influential data tools in the world
Build the next generation of AI-native data systems, including conversational agents, intelligent classification, automated governance, and more

Every team member receives an ownership stake in the company. When we grow, you grow with us.

All roles are remote unless otherwise specified in the job description. Review the job description to confirm if the role you are interested in is remote or hybrid.

Home office, coworking space, or something in between? We support your ideal setup. You’ll receive a monthly coworking stipend to use whenever you need a change of pace or in-person collaboration time.

Your well-being matters. We cover 99% of medical, dental, and vision premiums employees, and 65% for dependents.

We offer FSAs to help cover planned or unexpected healthcare costs. You can also opt into a Dependent Care FSA to support family needs.

Through Carrot Fertility, we provide inclusive fertility benefits and family-forming support. All U.S. employees have access, regardless of age, gender identity, or family structure.

We trust you to take the time you need. Our unlimited PTO and sick leave policy is designed for flexibility, rest, and real life.

 

Location & Eligibility

Where is the job
Worldwide
Fully remote, anywhere in the world
Who can apply
Same as job location

Listing Details

Posted
August 20, 2026
First seen
August 21, 2026
Last seen
August 21, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
68%
Scored at
August 21, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
DataHub
DataHub
greenhouse
Employees
5
Founded
2017
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

DataHubSenior Software Engineer or Staff Engineer- Ingestion Framework