Helia HR

Guide

Goals and OKRs for engineering teams (that people actually use)

Updated 2026-07-24 · For engineering managers and team leads at 5–200-person IT teams

OKRs, KPIs, SMART goals, metrics: one word per job

These four words get used interchangeably, and it makes goal-setting muddier than it needs to be. They name four distinct jobs:

  • A metric is just a number you can track. Deploy frequency, p95 latency, revenue, open bug count. A metric has no opinion — it is neither good nor bad until you decide what you want from it.
  • A KPI is a metric you have promoted to "watch this continuously". It is a health indicator with no end date: uptime, billable utilization, churn. You never "finish" a KPI — you keep it in a healthy range, quarter after quarter.
  • An OKR is a time-boxed goal to change something. An Objective (where you want to be by the end of the quarter, stated in plain, motivating language) plus Key Results (the measurable evidence that you got there). OKRs are for the two or three things you are deliberately trying to move this quarter — not the whole job.
  • A SMART goal is a quality test, not a framework. Specific, Measurable, Achievable, Relevant, Time-bound — a checklist you run over any single goal (including a key result) to catch the vague ones. SMART does not compete with OKRs; you apply it to them.

The rule of thumb that keeps them straight: if it has no finish line and you just want it to stay healthy, it is a KPI; if it is a change you want to achieve by a date, it is an OKR; SMART is how you check that either one is written well. "Keep uptime above 99.9%" is a KPI (or an SLO) — not an OKR, no matter how you format it. "Get p95 checkout latency under 300 ms this quarter" is an OKR, because there is a change and a deadline.

In a small engineering org you need all three, lightly: a short dashboard of KPIs you glance at (is the business healthy?), one or two OKRs a quarter (what are we deliberately changing?), and the SMART test as a five-second sanity check on each key result. Reach for more machinery than that and you have built process for its own sake.

Output vs outcome: the mistake that kills engineering OKRs

If engineering OKRs fail for one reason above all others, it is this: they get written as a to-do list. "Ship feature X. Migrate to service Y. Release the mobile app." Every item is real work and every item is checkable — and the whole set is output, a backlog wearing an OKR costume. You can complete all of it, grade yourself a perfect score, and change nothing anyone outside the team would notice.

The distinction is the whole game:

  • Output is what you build — the features, migrations, tests and releases you produce.
  • Outcome is what changes because you built it — faster, more reliable, higher-converting, less error-prone; something a user or the business actually feels.

Output key results feel safe because they are entirely under your control, and that is exactly the problem: they let you pre-commit to a solution before you have learned whether it works, and they reward motion over result. Outcome key results are scarier — you can work hard and still miss — which is what makes them honest.

The rewrites, all real engineering shapes:

  • Bad: "Ship the new onboarding flow." Good: "Raise new-workspace activation (setup completed within 24 hours) from its current level to a clearly higher one." The flow is a bet; the key result is the result the bet is supposed to produce.
  • Bad: "Migrate 12 services to the new CI pipeline." Good: "Cut median CI run time from around 18 to under 6 minutes, and reduce failed deploys from roughly 1 in 10 to under 1 in 30." Nobody outside the team cares that 12 services moved; they care that shipping got faster and safer.
  • Bad: "Write 200 integration tests." Good: "Halve the defects that escape to production per release." Two hundred tests that do not move the escape rate were the wrong two hundred.
  • Bad: "Hold three architecture reviews." Good: "Bring the payments service change-failure rate down to a stated target." The review is a tactic; the failure rate is the point.

The tell is a single question: could you complete this key result by working hard and still have improved nothing? If yes, it is output — keep asking "and then what?" until you land on something a user or the business would feel.

One honest caveat. Some quarters really are delivery quarters — a contracted client deadline, a compliance date, a launch that simply has to happen. A milestone key result is legitimate then. Just name it as a delivery commitment instead of dressing a deadline up as an aspiration, and keep at least one outcome key result beside it so you notice if the thing shipped and still did not matter.

How to write one: an objective and two to four key results

The anatomy is small on purpose.

  • The Objective is qualitative and memorable — where you want the team to be by the end of the quarter, phrased so it survives being said out loud without a spreadsheet. "Customers stop noticing us because nothing breaks" beats "Improve reliability."
  • Two to four Key Results are the measurable proof. Each needs a baseline and a target — a from and a to. If you cannot state today's number, your first key result is to instrument it; a key result with no baseline is a wish.

Four worked examples across the areas engineering OKRs usually live — reliability, delivery speed, quality, and team growth — each in a weak and a strong version. The numbers are illustrative; use your own baselines.

Reliability

Weak. Objective: "Improve reliability." Key results: "set up more monitoring", "reduce flaky tests." Both key results are output, and the objective could mean almost anything.

Strong. Objective: "The platform is boring — customers stop noticing it is there." Key results: p95 API latency from around 800 ms to under 300 ms; unplanned downtime from a few hours a month to under thirty minutes; the top five user journeys each covered by alerting with an owned runbook, from a starting point of none.

Delivery speed

Weak. Objective: "Ship faster." Key result: "release every sprint." That is a schedule, not an outcome.

Strong. Objective: "An idea reaches customers in days, not sprints." The four DORA delivery measures make a ready-made menu here — lead time for changes from around 9 days to under 3; deploy frequency from weekly to daily; change-failure rate held under a stated ceiling; time to restore service from hours to minutes. Pick the two that hurt most rather than chasing all four at once.

Quality

Weak. Objective: "Fewer bugs." Key result: "write more tests." Tests are output; more of them can happily coexist with the same bug rate.

Strong. Objective: "The team trusts main enough to deploy on a Friday." Key results: escaped defects per release cut by half; share of changes reverted within 48 hours held under a small target; flaky-test rate in the suite under a stated threshold.

Team growth

Weak. Objective: "Grow the team." It is ambiguous — headcount? skills? — and half of it is not the team's to control.

Strong. Objective: "Every engineer has a next step they are actively moving toward." Key results: every engineer has a current growth goal tied to the grade ladder; a new hire's time to first meaningful merged change from around 3 weeks to under one; at least two people with a documented promotion case in progress. Growth goals belong on a longer arc than a quarter — connect them to your career ladder rather than inventing quarterly targets for individual people.

Distilled to rules you can hold in your head: one to two objectives per team per quarter; two to four key results each; every key result has a from and a to; mix leading indicators (you can move them now) with lagging ones (they prove it worked); and if the objective needs a chart to be understood, rewrite it.

Cadence: quarterly goals, weekly check-ins, and why annual is too slow

Set OKRs quarterly. A quarter is long enough to move a real metric and short enough that a wrong bet costs you three months, not twelve. It also matches how engineering reality actually arrives — in reorgs, client changes, platform shifts and incidents that no annual plan survives.

Annual OKRs are for company strategy, not engineering execution. A team objective written in January is usually fiction by April; keeping it on the wall out of stubbornness only teaches everyone that OKRs are theatre. Let the year hold the vision and the quarter hold the goals.

Check in weekly or biweekly — 15 minutes, not a status meeting. Fold it into the top of a planning meeting you already have so it costs nothing new. Per key result, three things: the current number, a confidence call (on track, at risk, or off), and the one blocker. The entire purpose is to catch an "at risk" in week three, while you can still act — not to admire green bars.

Take a harder look mid-quarter, around week six. Not just "how are we tracking" but "is this still the right objective?" If reality moved — a client emergency, a shifted priority, something you learned — re-scope or drop an OKR openly. Adjusting because you learned something is the system working; refusing to adjust is the system failing.

Grade at the end, then talk about why. Score each key result — the classic 0.0-to-1.0 scale, or a plain red / amber / green — and then spend the real time on the retro: why did each land where it did, and what did you learn for next quarter? The grade is a conversation starter, not a report card, and (see the anti-patterns) it never feeds anyone's pay.

Aligning individual growth with team OKRs and reviews

Two kinds of goals get tangled together constantly, and separating them removes most of the confusion:

  • Team OKRs are what the team is trying to change this quarter — shared, outcome-focused, time-boxed.
  • Individual growth goals are how a person is trying to grow over quarters and years — personal, tied to skills and grade, on a much longer arc.

They connect, but they are not the same list, and you should resist turning them into one. In particular, do not cascade — you do not need every engineer to carry a personal mini-OKR derived from the team's. That is coordination machinery for hundreds of people; in a team of fifteen it is just paperwork. Team-level OKRs plus individual growth goals is enough.

The light version of alignment: every person should be able to point to how their work maps to a team key result — and, where it fits, a growth goal can ride on a team OKR. "You own the latency key result this quarter" can be exactly the end-to-end ownership evidence someone needs for their next grade. Anchor those growth goals in your career ladder, not in quarterly targets invented for individuals.

Reviews are where this most often goes wrong. OKR outcomes belong in a [performance review](/guides/performance-reviews-it-team) as context, never as the score. A review looks at how a person worked and grew across the whole period; an OKR grade is one input among many. The moment a team's OKR grade sets an individual's rating, target-setting turns defensive — people quietly pick goals they cannot miss — and you have broken OKRs as a planning tool to gain nothing. Keep the wall clean: OKRs plan the work, reviews assess the person, and the ladder carries pay.

Review cycles in Helia: periods, self and manager assessments, cycle status

Anti-patterns to avoid

Most OKR failures are one of a handful of predictable traps:

  • Too many OKRs. Three objectives with four key results each is twelve results nobody can hold in their head — and when everything is a priority, nothing is. One or two objectives, two to four key results each. If you cannot recite them from memory, you have too many.
  • Sandbagging. Setting targets you are certain to hit so the end-of-quarter grade looks green. Punishing it makes it worse; the cure is cultural — separate grades from consequences, and say out loud that a 0.6 on a real stretch beats a 1.0 on a safe bet.
  • OKRs as a stick in reviews. The single fastest way to kill honest goal-setting. If missing an OKR is dangerous, people will only ever set OKRs they cannot miss — and now the tool measures caution, not ambition.
  • Vanity metrics. Key results that go up and to the right without connecting to anything real — lines of code, story points, raw ticket counts, deploys for their own sake. A metric is vanity if it can improve while the thing you actually care about gets worse. Pair every count with an outcome or quality guardrail.
  • Copying Google verbatim into a 15-person shop. The 0.0-to-1.0 grading ritual, the moonshot language, the company-wide cascade — that machinery exists to coordinate thousands of engineers. Take the two ideas that travel (an inspiring objective, a few measurable results, reviewed quarterly) and leave the ceremony. Process borrowed at the wrong scale is how OKRs earned their bad reputation.
  • Set-and-forget. OKRs written at a quarterly offsite and never opened again until the next one. Without the weekly check-in they are not goals — they are a wish list with a nice heading.

A quarter in practice: a checklist for a lead with no HR or ops

If you are a team lead running this yourself, with no dedicated HR or operations support, the whole cycle fits in about two hours of setup and fifteen minutes a week. One quarter, start to finish:

  1. Before the quarter (60–90 minutes). Pick one or two objectives from what genuinely matters right now. Draft two to four key results each, every one with a from and a to; if you cannot name the baseline, your first key result is to measure it. Run each through the output-vs-outcome question. Write them somewhere the team sees them daily — not a doc that gets opened twice a year.
  2. Week 1: share and stress-test. Put the draft in front of the team and let them poke holes. A key result the team argued about is one they own; a mandate handed down is one they merely tolerate.
  3. Every week (15 minutes). At the top of a meeting you already hold, each key result gets its current number, a confidence call, and one blocker. Act on "at risk" the week it shows up, not at the end.
  4. Around week 6: mid-quarter honesty. Ask whether each objective is still the right one. Re-scope or drop openly if reality moved, and write down what you are learning — that is half the value.
  5. Final week: grade, then retro. Score each key result, then spend the real time on why it landed there and what it teaches for next quarter. Grades stay out of pay and reviews.
  6. Roll into the next quarter. Carry work that is unfinished but still important as a fresh key result — reworded, not merely extended — and retire whatever stopped mattering. A stale OKR dragged across three quarters is a signal, not a goal.

That is the entire system. Anything heavier at your size is the ceremony that makes engineers roll their eyes at the word "OKR" — and the point was never the ritual; it was a team that knows what it is trying to change, and can tell whether it did.

How Helia HR does this

Helia HR has goals and reviews built in, so the quarter's OKRs and a person's long-arc growth live next to the career ladder and the review — not in a spreadsheet nobody reopens:

  • Goals with measurable key results, grouped into periods. Set an objective, attach key results with a from and a to, and file it under a named period (like 2026-Q2) so the quarter's goals group together.
  • Progress that matches the check-in. Each goal carries an on-track / at-risk / off-track status, so the weekly 15-minute review is a glance, not a meeting.
  • Alignment without cascade paperwork. A goal can hang off a parent goal, so an individual's work maps to a team objective where it fits — without forcing a mini-OKR onto every engineer.
  • Kept out of the pay decision, on purpose. Goals feed the review cycle (self plus manager assessment) as context, while promotion and grade live on the separate career-path ladder — the exact separation this guide argues for.
  • A first draft when you're stuck. Helia AI can draft an individual growth plan from someone's grade criteria and history, which a manager then edits and owns.

Goals, reviews, and career paths are add-on packs on top of the base — switch on only the parts you actually run.

Career ladders in Helia: grades, criteria and promotion readiness — where growth goals live

FAQ

What is the difference between an OKR and a KPI?

A KPI is a metric you watch continuously to confirm something stays healthy — uptime, utilization, churn — with no end date. An OKR is a time-boxed goal to change something by the end of a quarter: an inspiring objective plus measurable key results. Rough rule: KPIs are the dashboard, OKRs are the roadmap. A key result can be "move this KPI from X to Y", but a KPI you only want to hold steady is not an OKR.

How many OKRs should a team have?

For a small engineering team, one to two objectives with two to four key results each, per quarter. If you cannot recite them from memory, you have too many — the entire point is focus, and a dozen "priorities" is the same as none.

Should OKRs affect performance reviews?

As context, yes; as a score, no. The moment an OKR grade sets someone's rating, people set targets they cannot miss and honest planning dies. Let the review look at how a person worked and grew, with OKR outcomes as one input among several, and keep pay decisions on the career ladder.

How often should we check in on OKRs?

Weekly or biweekly, about fifteen minutes, folded into a meeting you already have: current value, confidence, and one blocker per key result. Add a harder mid-quarter review around week six to re-scope if reality has moved. The check-in exists to catch an at-risk key result early — not to admire progress.

What is a good OKR grade?

On the classic 0.0-to-1.0 scale, hitting everything at 1.0 usually means the targets were too safe; a genuine stretch that lands around 0.6 to 0.7 is often the healthier result. Grades are a learning signal for the retro, not a report card — the question that matters is "what did we learn?", not "did we hit 100%?".

Should we use OKRs or SMART goals?

They are not rivals. OKRs are the quarterly structure — one objective and a few measurable results; SMART (specific, measurable, achievable, relevant, time-bound) is a checklist you apply to each key result to catch the vague ones. Use OKRs to decide what matters this quarter and the SMART test to make sure each result is written well.

Run HR and delivery ops in one system

Helia HR combines the HR basics with the capacity matrix, bench view, timesheets and client invoicing IT services teams actually run on. Start free, no card. GDPR-grade security, role-gated PII, audit-logged access.