Data Engineer
WHY THIS ROLE EXISTS
Every city, county, and school district buys things. Roads, software, cleaning services, IT infrastructure. The total is somewhere north of two trillion dollars a year. And almost all of it moves through procurement processes designed for a different era: slow; paper-heavy; opaque and exhausting for everyone involved.
Civic Marketplace was built to fix that. We're a modern, data-driven platform where government agencies discover, evaluate, and engage suppliers. Where businesses, especially smaller and growing ones, can actually find and win public sector work without needing a dedicated contracts team to navigate the maze.
We're past the point of proving this works. Agencies are live on the platform, real money moves through it, and we're now combining that marketplace infrastructure with AI in ways that could genuinely transform how procurement works. Not just incrementally but structurally.
None of it works without data that can be trusted. Procurement is a data problem before it is a software problem: who the suppliers actually are, what they can genuinely deliver, which contracts an agency is already entitled to buy from, what this thing cost the county next door. That information exists, scattered across thousands of agency portals, state registries and PDF attachments, in no agreed format, and nobody has assembled it properly. Whoever does gets to define how public money is spent for the next decade. That's this role.
THE PROBLEM YOU'D OWN
The hard part isn't moving data from one place to another. Any engineer here will tell you the pipelines are the easy half.
The hard part is that public procurement has no shared vocabulary. The same supplier turns up as four different legal entities across three registries, and two of the spellings are wrong. Commodity codes are applied inconsistently, or not at all. A cooperative contract one agency can buy from today is invisible to the agency next door because it was published as a PDF on a portal with no API. Every source you touch is incomplete, inconsistent, or both, and most of it is public record, so you can't quietly correct it. You have to model the mess honestly.
And the bar just moved. We recently launched an MCP integration that lets agencies request quotes through Claude, GPT and Copilot. When an agent answers a procurement question, a wrong answer doesn't read like a bug, it reads like advice, and in public spending, bad advice ends up in a council meeting. That puts weight on freshness, lineage and provenance that most product data layers never carry. As we build out our agentic procurement capabilities, the reliability of that data, and how we prepare it for retrieval, becomes our most critical engineering challenge.
So this is an entity resolution and data trust problem dressed as a pipeline problem, and it's yours to solve. You'd work out what is actually wrong with the data, which is rarely what everyone assumes, then build the fix so it holds for every source we add next. Not a script per source held together by whoever wrote it.
ABOUT THE ROLE
You would be our first dedicated data engineering hire, reporting to Mikey Mo https://www.civicmarketplace.com/team/mikey-mo, our Head of Engineering. Data work today sits with the product engineering team and gets done alongside shipping features. It works, but it belongs to whoever last touched it, and that isn't a foundation for what comes next.
To give you a sense of the gap: awarded quotes that slip through the platform get manually caught by our commercial team and hand-entered into HubSpot instead. Supplier onboarding data follows the same pattern. So as and when the platform and HubSpot become out of sync, it makes it difficult to say with confidence which one is right.
This is a build-and-do role. You'd collaborate closely with our Head of Engineering to shape the architecture, and you'd also write the pipelines, carry the pager for them, and go and read the raw source when a number looks wrong. If you're looking for a role where you set direction in isolation and hand implementation to other people, this isn't it. If you want substantial responsibility and scope as a function lead alongside our Head of Engineering while still enjoying the craft yourself, it very much is.
You would not be starting from nothing. There's a live platform with real transactions running through it, a product engineering team who know where the bodies are buried, and access most people in this field would have to scrape for: Civic Marketplace is a member of the NIGP Business Council, with real partnerships across councils of governments. Most people doing this job spend their first year getting the data access. You'd start with the door already open.
WHAT YOU'D OWN
Five things, and the freedom to decide how.
1. Trust in the data. The north star. A golden, trusted dataset that powers Civic Marketplace's analytics, so when someone asks how many quotes were awarded last quarter, or how much GMV we've captured, there's one number, and it's right. You'd own the diagnosis and the fix.
2. Ingestion. Pipelines that pull supplier, agency, solicitation and contract data out of sources never designed to be read by anything but a person. Reliable, observable, and cheap enough to add the next source without a meeting about whether it's worth it.
3. The canonical model. One supplier, one record, across every spelling, trading name, subsidiary and registry identifier. One shared definition of a contract vehicle, a commodity, an agency, that product, sales and customer success all use and none of them argue with. This is the unglamorous part that makes everything downstream possible.
4. The system underneath. Testing, lineage and freshness, so when the platform tells an agency something we know where it came from and when. And self-serve data for product, customer success and analytics, so nobody queues behind an engineer for a number. Everything you build should work for the next source without you in the loop. This is where AI earns its place, turning what is currently hand-checked into something repeatable.
5. Agentic data infrastructure. As we scale our use of AI-native procurement, you will own the data layer that powers our agentic workflows. This goes beyond standard retrieval. You will design ingestion and storage strategies, including optimizing vector pipelines like Pinecone for RAG, to ensure our agents have the high-fidelity, context-aware data necessary to perform reliably in production.
You'd work closely with product engineering, customer success and our Growth Lead. Application feature work stays with the product engineering team, so you're not competing for that ground. You own the layer they build on top of.
HOW WE USE AI
We use Claude daily, across content pipelines, community response and internal knowledge, and increasingly inside the product itself. That's real, not a line in a job ad.
For this role it cuts two ways. There's how you work: source profiling, schema mapping, test generation, the documentation that otherwise never gets written. Most of that is now automatable, and we expect you to automate it. And there's what you build: the data layer that decides whether agentic procurement is trustworthy or embarrassing. The second is the harder and more interesting problem.
What we care about is not whether you can use it. Everyone says they can. It's whether you reach for it to make something repeatable. The difference between someone who writes pipelines faster and someone who builds the tooling that means the next hundred sources don't each need hand-holding is the difference we're hiring for.
If you've built data quality tooling, evaluation harnesses, retrieval pipelines or automated documentation from scratch, tell us about them. If AI has changed how you think about the work and not just how fast you produce it, we especially want to hear that.
Our stack: Snowflake, Postgres, dbt, Fivetran, HubSpot, PostHog, Metabase.
WHAT YOU BRING
YOU'LL NEED:
- Data engineering experience in a startup or scale-up, where you've built the platform rather than inherited one
- Strong SQL and Python, and the judgement to know when the warehouse is the wrong place to solve a problem
- A data model you designed that other people had to live with, including the parts you would do differently now
- Real experience of messy external sources: half-documented APIs, flat-file drops, scraped pages, PDFs, and data you neither control nor can correct
- Entity resolution or record linkage experience, or clear evidence you'd be good at it. This is the centre of the job, not an edge case
- A commercial head. You can tell the difference between a data problem that's blocking revenue and one that's merely interesting, and you sequence accordingly
- The instinct to find out why a number is wrong before proposing a fix. Sometimes it's the pipeline, sometimes the model, sometimes the source was always like that, and those need different answers
- The instinct to build systems that scale rather than pipelines that run once
- Genuine AI fluency, in the sense above
- The determination to stay positive through the ups and downs of a fast-moving startup, and to find a path through problems rather than wait for conditions to be right
- Comfort operating remotely across timezones with real autonomy and not much oversight
- Curiosity about why public sector data is different. You don't need to have done it. You do need to want to understand it
YOU'LL STAND OUT IF YOU HAVE:
- Govtech, public sector or civic tech background, or experience working with public records and open data
- Familiarity with supplier and procurement data: SAM.gov http://SAM.gov and UEI, DUNS, NAICS or UNSPSC, cooperative purchasing, COG or NIGP
- Experience with search and retrieval, whether OpenSearch, Elasticsearch or vector pipelines (like Pinecone or Milvus) feeding an LLM product.
- Bilingual English and Spanish, especially valuable for supplier data quality and onboarding
- Experience being the first data hire on a team that had been doing it themselves
WHY YOU'LL LOVE THIS ROLE
The problem is genuinely important. Public procurement touches every road, school and hospital. Getting it right matters in ways most B2B SaaS doesn't. And you'd be widening access for small and local businesses, which is the part of this that keeps us up at night in a good way.
The dataset doesn't exist yet. Most data engineering jobs are being the fifth person to tidy the same warehouse. This one is assembling something nobody has assembled properly: a clean, current picture of who supplies the public sector and what public money actually buys.
The scope is unusual. First dedicated data hire, reporting to Mikey, taking substantial responsibility for the function and co-designing the architecture decisions that come with it.
You'd own a real piece of it. Equity is part of the package here, and we mean it as more than a line in an offer letter. We're looking for someone who wants to build something they have a stake in, not just a job with a good salary attached.
The AI opportunity is real. We're not bolting AI onto existing workflows. We're rethinking what's possible when procurement becomes genuinely agentic, and the data layer is the part that decides whether any of it can be trusted.
People who care. Small team, direct access to founders, genuine investment in doing things well. We move fast, but we don't move sloppy.
HOW WE WORK
Build bridges to help customers win. We are obsessively focused on helping both agencies and suppliers succeed. That's not a value statement. It's the job.
High velocity, high ownership. We move quickly, make decisions, and take responsibility for outcomes. There's no one to hand things off to and no one to hide behind.
In the arena. We stay close to our users. We learn from them directly. We build based on what we see, not what we assume.
Learning quotient. Rapid iteration isn't just a process. It's a mindset. We'd rather be wrong fast and right eventually than slow and cautious throughout.
And because we work in public procurement, how we do things matters as much as what we achieve. We hold ourselves to the standards our agencies are held to.
THE INTERVIEW PROCESS
All interviews are via video conferencing by default.
We'll ask you to walk us through something you've built and the decisions you'd make differently now. We'll also set you a short exercise on genuinely messy source data rather than an abstract algorithm puzzle, because reasoning about real data is the job. We'll talk about how you think, not just what you've done. And we'll tell you everything we know about where Civic Marketplace is going, because we want you to be choosing us as much as we're choosing you.
WHAT WE OFFER
- Competitive salary and early-stage equity
- Comprehensive medical, dental, and vision insurance
- Flexible PTO
- Remote-first, with real flexibility across timezones (Remote, USA; London, UK)
- Full AI tool stack: Claude Pro, HubSpot, Make, Notion, and more
- Regular team offsites, including international meet-ups (ask us about Reykjavik!)
- Direct access to the founding team and a front-row seat to building something that matters
Civic Marketplace, Inc is an equal opportunity employer. We actively encourage applications from candidates of all backgrounds, identities, and experiences.
Sourced from a public career listing. Jobverse is an aggregator, not the employer.