~ / about

Fourteen engineers. No account managers.

GridWorks was founded in 2019 by three engineers who had spent a decade being the people called at 3am. We stayed deliberately small: everyone who works on your systems has operated systems at scale themselves, and the person in the sales conversation is the person who will be in your repositories.

A long, low-lit open-plan engineering floor with rows of desks and monitors.
Bristol, 2019 · the room we started in / now remote-first across 5 timezones
Founded
2019
Bristol, then everywhere
Engineers
14
all staff-level or above
Engagements delivered
96
across 38 companies
Upstream commits merged
1,740+
CNCF projects, since 2019
Engineering principles

Eight things we will argue about

These are not values on a wall. They are the positions we take in design reviews, and the reason some engagements do not happen.

PRINCIPLE 01

Boring technology, deliberately chosen

Every novel component in your stack is an on-call burden and a hiring constraint. We spend innovation tokens where they create differentiation for your product, and nowhere else. Postgres until Postgres genuinely stops working.

PRINCIPLE 02

If it isn't measured, it isn't done

We do not accept "it feels faster" as an outcome. Every engagement opens by establishing a baseline — latency histograms, change failure rate, cost per request — and closes by comparing against it. Including when the comparison is unflattering.

PRINCIPLE 03

Reliability is a budget, not an absolute

Chasing five nines on a service where users would not notice four is an expensive way to slow your roadmap down. We help you pick a target, defend it, and spend the remaining budget on shipping.

PRINCIPLE 04

The platform serves the product team

A golden path that engineers route around is a failed platform, no matter how elegant the abstraction. We measure adoption, we ask why people opt out, and we treat the developer as the customer.

PRINCIPLE 05

Automate the recovery, not just the deploy

Anyone can automate the happy path. The value is in the rollback that runs without a human, the failover that has been drilled this quarter, and the restore that someone has actually tested against production-sized data.

PRINCIPLE 06

No lock-in, including to us

We build on upstream open source, in your accounts, in your repositories. There is no GridWorks agent, no proprietary controller, no licence. If we are still essential after twelve months, we have done the job badly.

PRINCIPLE 07

Incidents are a systems problem

Nobody has ever been fired at a GridWorks post-incident review, because "human error" is the beginning of an investigation, not the end of one. If a single mistake could take production down, the system permitted it.

PRINCIPLE 08

Say no to the wrong engagement

We have turned down work where the real problem was organisational, where the timeline was fiction, or where the client wanted a rubber stamp on a decision already made. We will tell you that in the first call rather than the third invoice.

The team

Who actually shows up

Six of the fourteen; the rest are on client engagements and prefer not to be listed. Every principal has run production at a scale where the failure modes get interesting.

Portrait of Dr. Anja Vosloo

Dr. Anja Vosloo

Founder · Principal, Platform

Spent six years on edge infrastructure at a global CDN, where she led the migration of the anycast routing layer off a bespoke control plane. PhD in distributed consensus; still maintains a Raft implementation nobody asked for.

  • Multi-region
  • Edge / anycast
  • Consensus
Portrait of Marcus Ilori

Marcus Ilori

Founder · Principal, SRE

Nine years of SRE across a hyperscaler and a UK challenger bank, the second of which taught him that regulators and error budgets are compatible if you write things down. Wrote the incident command handbook we give clients.

  • SLO design
  • Incident command
  • Regulated ops
Portrait of Priya Raghunathan

Priya Raghunathan

Principal, Delivery Engineering

Built the deployment platform at a commerce company through the years it went from 80 to 900 engineers. Has strong opinions about merge queues and a documented grudge against long-lived feature branches.

  • CI/CD at scale
  • Monorepo
  • Progressive delivery
Portrait of Tomás Ferreira

Tomás Ferreira

Staff Engineer, FinOps

Former capacity planner on a public cloud compute team, so he knows what the pricing page does not tell you. Built the unit-economics model we use, and once found a $340k/year saving in a single misconfigured storage lifecycle rule.

  • Unit economics
  • Commitments
  • Capacity modelling
Portrait of Ines Marchetti

Ines Marchetti

Staff Engineer, Observability

OpenTelemetry collector contributor and maintainer of two receivers. Came from an APM vendor's backend team, which is why she is unusually direct about what telemetry actually costs to store and query.

  • OpenTelemetry
  • eBPF
  • Query cost
Portrait of Kwame Osei

Kwame Osei

Staff Engineer, Supply Chain Security

Worked on artifact signing and provenance tooling before it had a marketing category. Leads our SLSA hardening work and reviews every policy pack before it reaches a client cluster.

  • SLSA / provenance
  • OPA · Kyverno
  • Zero standing access
Two engineers reading the same laptop screen during a paired working session. Paired working session · half a day, paid

How we hire

No take-home exercises and no algorithm rounds. Candidates walk us through a system they have operated and an incident they got wrong, and then pair with two of us on a real client-shaped problem for half a day, paid. We have made eleven hires this way and lost one.

Minimum bar
7+ years operating production systems
On-call
Everyone, including founders
Location
Remote, UTC−5 to UTC+5:30
Utilisation target
65% — the rest is upstream work and research
Upstream

What we give back

Every engineer gets one day a week for upstream work. It is not altruism — patching a bug in the collector is cheaper than maintaining a fork of it across nine clients.

Project Area Contribution Status
opentelemetry-collector-contrib Telemetry Two receivers maintained, 340+ commitsplus the tail-sampling memory regression fix in v0.98 maintainer
karpenter Compute Consolidation edge cases with PDB interaction6 merged PRs, 2 open RFCs contributor
argo-rollouts Delivery Prometheus analysis provider — exemplar supportand a fix for aborted-rollout metric leakage contributor
crossplane-contrib/provider-aws Substrate RDS composition defaults, PITR and CMK wiringreduced a 200-line claim to 12 contributor
kyverno Policy Image verification policy library for cosign keylessnow part of the upstream sample set contributor
gw/slo-generator Our own Generates multi-window burn-rate rules from a service manifestApache 2.0, 2.1k stars maintainer
gw/tf-module-library Our own The 41 Terraform modules we hand to clients, publishedApache 2.0, terratest coverage 84% maintainer
↔ scroll table horizontally
gw oss stats --year 2026
commits merged      412
projects touched    19
issues triaged      288
CVEs reported       3
talks given         7
──────────────────────────
maintainer seats    3
forks maintained    0
  ↑ this is the metric
    we actually care
    about
Fit

When to call us

  • Your deploys need a change window and a person watching a dashboard
  • You are about to sign a multi-year cloud commitment and cannot defend the baseline
  • The same three engineers are the only ones who can safely touch production
  • An audit or customer contract has introduced a resilience requirement you cannot evidence
  • Cloud spend is growing faster than revenue and nobody can attribute it
  • You are hiring a platform team and want the foundation laid before they arrive
Anti-fit

When not to

  • You want bodies on seats — we are not a staffing firm and will be expensive at it
  • The decision is already made and you need external validation for it
  • Nobody internally will own the platform after we leave
  • The real problem is that two teams will not talk to each other
  • You need a managed service with a support portal, not an engineering engagement
  • The timeline was set before the scope, and cannot move
gw review --schedule

Talk to an engineer, not a funnel

The first conversation is with whoever would lead your engagement. If we are not the right fit we will say so on that call and, where we can, point you at someone who is.