Engineering Manager, Cloud Monitoring Services Platform
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About the Role:
Crusoe builds cloud infrastructure for AI workloads. Cloud Monitoring Services owns observability across Crusoe Cloud: metrics, logs, alerting, and the telemetry agent that runs on every node in the fleet. Our customers run large, demanding AI training and inference workloads, and they depend on us for a clear, trustworthy view of what their infrastructure is doing.
We are hiring an Engineering Manager to lead the Platform team: the time series and log storage systems, and the query layer that serves every dashboard, API call, and investigation on top of them. This is a first-line management role reporting to the Engineering Manager for Cloud Monitoring Services. You will take direct people management responsibility for a team of 4 to 6 engineers, growing, and own how telemetry is stored, retained, and read back at fleet scale.
This is a people-first leadership role with real delivery stakes. You will bring the technical depth to guide hard calls and earn your team's trust, but your success is measured through what your team accomplishes. Query latency, retention, and storage cost are live tradeoffs your team will be making continuously, and customers feel all three directly. The role sits alongside peer managers who own telemetry collection and the ingestion pipeline, and together you run a platform that has to work end to end. This is a full-time position.
What You'll Be Working On:
- Grow and develop your team. Manage 4 to 6 engineers directly: 1:1s, career growth, performance, and team health, with the team growing over time.
- Own storage and query. Be accountable for the time series and log storage systems and the query layer on top of them, including how they behave under load and how much they cost to run.
- Own delivery. Plan and sequence work across a roadmap that mixes customer-facing feature work with storage and query infrastructure, and make honest calls early when a plan is at risk.
- Set the technical direction for your area. Partner closely with the Staff engineers who own technical direction across the platform, so they can focus on engineering rather than absorbing delivery and handoff work alone. Ask the hard questions and make sound tradeoff calls alongside your engineers.
- Keep the operational bar high. Query is on the critical path when something is wrong in the fleet, so availability, query performance, and correctness of what comes back are core to the job, not afterthoughts.
- Own the team's oncall rotation and the operational health of the storage and query stack.
- Manage cost and scale together. Retention policy, downsampling, cardinality, and storage tiering are product decisions as much as engineering ones, and you will be in those conversations.
- Coordinate across team boundaries. Your team consumes what the collection and ingestion teams produce and serves internal platform teams as well as external customers, so sequencing dependencies and escalating early is a core part of the job.
- Shape the team over time. Work with recruiting on sourcing, run a high-quality interview loop, close strong candidates, and onboard them well.
- Collaborate across functions. Partner with product, neighboring infrastructure teams, peer managers, and leadership to keep priorities aligned as the observability product expands.
What You'll Bring to the Team:
We know great engineers and leaders come from many different paths. If you're excited about this work but don't match every point below, we'd still love to hear from you.
- You care about people. You want the engineers around you to grow, you give feedback that's both direct and kind, and you measure your own success through your team's.
- Technical depth. 5+ years of hands-on engineering experience, ideally in backend or distributed systems: databases, storage engines, query engines, streaming pipelines, or other data-intensive services (Go, Rust, Java, or C++ environments), and hands-on familiarity with Kubernetes. This role stays close to the technical work.
- Experience managing engineers. 2+ years managing software engineers directly, including owning performance cycles and career conversations. This team has an active roadmap and live customer commitments, so we are looking for someone who has run a team through delivery before.
- Experience coaching senior and staff-level engineers, and comfort partnering with strong technical leads rather than competing with them.
- A track record of shipping. You've delivered multi-phase projects against fixed deadlines, and you know how to keep infrastructure work moving alongside feature work rather than letting one crowd out the other.
- Operational judgment. You've run a team that owns a system customers depend on during an incident, and you know what it takes to keep that system trustworthy.
- Strong communication and judgment. You can align people across functions, explain tradeoffs to technical and non-technical partners alike, and give your team clarity about what matters and why.
Bonus Points
- Experience with time series databases or columnar storage at scale: Prometheus, Thanos, Mimir, VictoriaMetrics, InfluxDB, ClickHouse, or similar
- Experience with query engines and query cost control: planning, pushdown, caching, concurrency limits, and protecting a cluster from expensive queries
- Experience with retention, compaction, downsampling, and storage tiering for large telemetry or event datasets
- Familiarity with log storage and search systems: Loki, Elasticsearch, OpenSearch, or similar
- Observability domain experience: OpenTelemetry, Grafana, metrics and log pipelines, cardinality management
- Experience owning multi-tenant systems with per-tenant isolation, quotas, and fairness
- Experience running software in GPU or accelerated computing environments
- Experience managing or scaling a team through growth
Benefits:
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location
Compensation Range
Compensation will be paid in the range of up to $215,000 -$260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.
Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
Sourced from a public career listing. Jobverse is an aggregator, not the employer.