Skip to content

Grafana

AGPLv3

Know what your systems are doing before your customers tell you.

Practice led by Callum Reid, Practice Director, Observability

What Grafana is

Grafana with Prometheus for metrics and Loki for logs, deployed and secured against your existing estate. Not a monitoring product you subscribe to — an observability stack you run, with your telemetry staying in your infrastructure and no per-host or per-metric billing to design around.

Who it suits

  • Teams whose first warning of an outage is a phone call from a customer
  • Organisations paying per-host or per-GB to a SaaS monitoring vendor and finding the bill scales faster than the estate
  • Anyone with telemetry they cannot send to a third party for regulatory or contractual reasons
  • Operations teams inheriting an estate they did not build and cannot currently see into

And who it does not

If you have fewer than a dozen hosts and no compliance constraint, a hosted monitoring service is probably cheaper and simpler than anything we would build you. We have said so and lost the work.

What we do with it

Our Grafana work

Configuration is the project. The install is an afternoon.

The stack, deployed and secured

Grafana, Prometheus and Loki with retention sized to your actual data volume, authentication wired to your identity provider, and TLS throughout. Not a default install with the admin password changed.

Dashboards against your services

Built around what your systems actually do and what your team actually needs to see at three in the morning. Generic template dashboards look impressive in a demo and get ignored within a fortnight.

Alerts designed to be acted on

Every alert has an owner, a runbook and a threshold someone has defended. Routing to email, Mattermost, PagerDuty or Opsgenie. We will actively push back on alerts nobody will act on — an ignored alert trains people to ignore all of them.

Service level objectives

Where you want them: error budgets, burn-rate alerting, and a dashboard that shows whether you are meeting the commitments you have made to your own customers.

Shape of the work

How an engagement runs

  1. Estate review

    What exists, what emits telemetry, what is currently invisible.

  2. Deploy

    Stack stood up, secured, integrated with identity.

  3. Instrument

    Exporters, scrape configuration, log shipping.

  4. Dashboards and alerts

    Built with your team, not delivered to them.

  5. Handover

    Runbooks, and training so your team can extend it.

What it connects to

  • Prometheus, Loki, Tempo and Mimir
  • PagerDuty, Opsgenie and Mattermost
  • Kubernetes, Docker and bare metal
  • AWS, Azure and on-premise hypervisors
  • PostgreSQL, MySQL, SQL Server and Oracle exporters
  • SSO through SAML or OIDC

The honest bit

Grafana is AGPLv3, and that matters more than it does for the other platforms we deploy: if you modify Grafana itself and offer it over a network, the AGPL has obligations. In practice this almost never applies, because dashboards, alerts and plugins are configuration rather than modification. We will still walk your legal team through it if they ask, and we do not modify Grafana itself as a matter of policy.

The Grafana project ↗

Talk to the Grafana practice

A scoping conversation with someone who works on this platform every day, not an account manager.