Data Engineer
Build the pipelines behind datastore: decoding years of on-chain history into clean, versioned Parquet that researchers can trust.
- Location
- Toronto, Canada
- Type
- Full-time, hybrid
- Team
- Data
About CLR3
datastore sells something unusual: files, not API access. Customers download decoded on-chain history as Parquet and run their own queries. That only works if the data is actually right, which makes correctness, lineage and reproducibility the product.
You will build and run the pipelines that decode Solana and Hyperliquid history at scale: backfills over billions of rows, schema design, checksums, manifests and the quality checks that let a quant trust a file they did not produce themselves.
What you will do
- Design and run large-scale decoding and backfill pipelines
- Model typed schemas for instructions, events and state across dozens of protocols
- Build validation that catches bad data before customers do
- Keep versioning, checksums, manifests and lineage docs accurate on every delivery
- Tune storage layout and partitioning so files query fast in DuckDB and Polars
- Add new protocols and chains to the catalogue
- Support customer questions about schemas and coverage
What we are looking for
- Experience building production data pipelines at meaningful scale
- Strong SQL plus one of Python, Rust or Go
- Real familiarity with columnar formats, ideally Parquet, and query engines like DuckDB, Polars or Spark
- Care for data correctness: testing, validation and reconciliation
- Comfort owning pipelines in production, including when they break at night
- Able to work from our Toronto office part of the week
Nice to have
- Experience with blockchain data or other messy, high-volume event streams
- Familiarity with warehouse ecosystems your customers use, like Snowflake, BigQuery or ClickHouse
- Background in quantitative research support or backtesting infrastructure
- Experience with orchestration tools and with knowing when a cron job is enough
Who you are
- You think an unverified number is worse than no number
- You write pipelines you would be happy to debug at 2am, so they rarely need it
- You like schemas that make the next person's query obvious
- You get satisfaction from a backfill that reconciles to the last row
- You explain data problems in plain language
How we hire
- 1Intro call with an engineer, about 30 minutes
- 2Short take-home assignment working with a real decoded dataset
- 3Technical conversation about your assignment and pipelines you have run
- 4Conversation with the founders
- 5Offer
Apply
Email careers@clr3.org with a short note about yourself, a link to your GitHub or past work, and a CV if you have one. We read every application and reply to all of them.
Apply for this role