Principal Software Engineer - Observability & Telemetry Data
Quick Summary
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents.
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.
- Architect the Telemetry Data Platform: Own the end-to-end design for how telemetry is landed, modeled, and queried, including open table formats, partitioning and schema evolution strategy, separation of storage and compute, tiered retention, and the query interfaces engineers actually use. Make telemetry a durable, portable, Smartsheet-owned dataset rather than a vendor-locked byproduct.
- Own the Telemetry Data Model: Define the semantic conventions, shared business identifiers (user, org, plan, tenant), and schema standards that let any signal be correlated with any other, and drive their adoption across every service team at Smartsheet.
- Set OpenTelemetry Direction: Lead the migration to OTel-based instrumentation, defining collector architecture, context propagation, and sampling strategy (including tail-based sampling) so that engineers can move from a log line to a trace to a metric without losing the thread.
- Integrate AI and Agentic Telemetry with Databricks: Own the architecture connecting our observability platform to Databricks and MLflow, so that agentic and model telemetry (prompt, completion, tool and MCP calls, evaluation results) is captured un-sampled, stays useful under the input/output redaction our governance requires, and reconciles cleanly with the traces and metrics in our primary observability stack.
- Instrument the Data Platform Itself: Bring first-class observability to our data estate, including Databricks jobs, pipelines, and warehouses, with meaningful signals for freshness, data quality, lineage, and cost, so data reliability is measured with the same discipline as service reliability.
- Build the Analytics Layer on Telemetry: Turn telemetry into decision-grade analytics, covering reliability and incident metrics, telemetry cost and chargeback models, and adoption and coverage reporting that leadership can act on.
- Engineer Collection and Routing at Scale: Architect the high-volume collection and routing tier (FluentBit, Kinesis, and OTel collectors) that moves telemetry from every service to its destination across US, EU, AU, and GovCloud regions, and own the migration of these pipelines as we consolidate onto a unified backend.
- Own Telemetry Economics: Set the cost architecture for observability data, including ingest governance, cardinality control, and storage tiering, so that teams get the fidelity they need to debug without the spend pressure that causes them to under-instrument.
- Carry Architecture Across Org Boundaries: Partner with the Data Platform, AI Platform, and infrastructure organizations to align telemetry architecture with theirs, influence roadmaps you do not own, and represent observability in company-level platform and vendor decisions.
- Raise the Technical Bar: Lead design and code reviews, author the architecture decisions and standards others build against, and mentor senior and mid-level engineers on instrumentation, telemetry data modeling, and cost-aware design.
- Participate in a production support and on-call rotation, taking ownership of the most complex issue resolution and driving root-cause analysis that improves system resiliency.
- 10+ years of experience building and operating large-scale distributed systems, data platforms, or observability infrastructure, including time at Principal or Staff level.
- Deep Observability Expertise: Hands-on production ownership of metrics, logs, and distributed tracing at scale, including at least one major backend (Datadog, or comparable) and a working understanding of cardinality and cost mechanics.
- Telemetry Data Engineering: Demonstrated depth in large-scale data architecture, including open table formats (Delta Lake, Apache Iceberg), Spark or comparable distributed processing, streaming ingestion, partitioning and schema evolution, and query performance and cost tuning over very large datasets.
- Databricks Depth: Practical experience with the Databricks platform (jobs, clusters, Unity Catalog, Delta) and with MLflow for model and agent telemetry.
- OpenTelemetry Depth: Practical experience with OTel collectors, semantic conventions, context propagation, and sampling strategy, including tail-based sampling.
- Pipeline Engineering: Experience with high-volume log and telemetry pipelines, including FluentBit or Fluentd, streaming transport such as Kinesis or Kafka, and search backend index and mapping design.
- Advanced AWS & Kubernetes Expertise: EKS, ECS Fargate, EC2, Lambda, and CloudWatch in production.
- 10+ years of programming experience with modern languages such as Go, Java, Python, or Scala, and strong SQL.
- Infrastructure as Code: Terraform, and GitOps workflows such as Flux or ArgoCD.
- Architectural Influence at Scale: A track record of setting technical direction that multiple teams and organizations adopted, including the written artifacts (architecture decisions, standards, RFCs) that made it durable after you moved on.
- Strong incident response instincts, with experience improving mean-time-to-resolution through better instrumentation and better data rather than more heroics.
- A degree in Computer Science, Engineering, or a related field, or equivalent practical industry experience.
- Legally eligible to work in the U.S. on an ongoing basis
Nice to Have
~1 min read- Experience instrumenting LLM or agentic systems, including OTel GenAI semantic conventions and tracing agent tool-call workflows.
- Data reliability engineering practice: freshness, quality, and lineage SLOs for production data pipelines
- Telemetry cost engineering or FinOps at scale.
- Experience in regulated environments (FedRAMP, GovCloud) and designing telemetry that remains useful under redaction.
- Prometheus, Grafana, and Alertmanager.
- Snowflake, or experience operating across more than one lakehouse or warehouse platform.
- Service catalog or internal developer platform work (Backstage or similar).
What We Offer
~1 min readAt Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.
Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information.
If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.
#LI-Remote
Location & Eligibility
Listing Details
- Posted
- September 1, 2026
- First seen
- September 2, 2026
- Last seen
- September 2, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 79%
- Scored at
- September 2, 2026
Signal breakdown

The foundation for managing projects, programs, and processes that scale.
View company profilePlease let Smartsheet know you found this job on Jobera.
3 other jobs at Smartsheet
View all →Explore open roles at Smartsheet.
Similar Software Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.