Cycle Time in Software Development: Definition, Formula and Benchmarks
By Sebastiaan Jansen · · 13 min read
Cycle time is the elapsed time from the moment a team starts a piece of work until that work is finished. It answers one narrow question: once we said yes and began, how long did delivery take? A stable number is something you can plan around. When the values are spread all over the calendar, the average is hiding a queue.
What is cycle time
Two teams can use the same phrase and mean different clocks. One starts when a ticket moves to "In progress." Another starts when a branch is created, and a third when code is merged. All three are defensible, and none of them is the official one. Pick one clock, compute it the same way every week, and read it honestly.
A usable cycle time definition names three things: the start event, the end event, and which items go into the sample. Leave one out and you have an impression, not a measurement. "We feel faster" doesn't count. A finish timestamp minus a start timestamp does.
The cycle time meaning is practical: how long active delivery takes. It says nothing about how long the idea existed or how many items you finished. A feature can sit in a backlog for a quarter and then take two days to build. Those two days are this metric. The quarter belongs to the request clock, which is separated from the work clock below.
What the clock is counting
The clock counts calendar time between two events you chose in advance. Effort hours, story points, and how hard the work felt don't come into it. A ticket that starts Monday at 09:00 and finishes Wednesday at 09:00 took 48 hours, even if someone spent only six of those hours writing code.
That gap is the reason to measure at all. A ticket often waits on review, a test environment, a product answer, or a deploy window. Effort can look fine while customers wait, and elapsed time is what shows the wait.
What it is not
It is not the age of the oldest open ticket. Age is a snapshot of unfinished work, while this metric describes finished work.
It is also not velocity or utilization. Two teams can finish the same number of story points and still differ by weeks of waiting. A busy team can take three weeks on a small change because too many items are open at once. And if you start the clock at code review to hit a target, you delete the development wait from the record. Pick the interval you want to manage, then leave it alone long enough to see a trend.
Where the clock should start and stop
Write the policy in one paragraph and put it where the team can see it.
- Start at the first status that means a person is now responsible for finishing the item. For many product teams that is "In progress." For a platform team it might be "Accepted," if that is when the clock starts to matter to stakeholders.
- Stop at the status that means a customer, or the next team, can use the result. "Merged" is a stop only if merge equals release. If you batch deploys on Thursdays, Thursday is the stop. Otherwise you are measuring coding time and calling it delivery.
- Say which items count. Production defects, planned features, and spikes often behave differently, so either segment them or keep one pool and say so. Mix a five-minute copy fix with a three-week migration and the median goes green while releases slip.
- Decide whether weekends, nights, and holidays count. Calendar time matches what the customer experiences; business hours answer staffing questions. Default to calendar time.
Statuses that ruin the sample
A few board habits make the series useless:
- Items that skip the start column never show up, so the fastest work vanishes and the metric looks worse than it is.
- Items reopened after "Done" need a new interval, or a rule that the original finish stands and a new ticket tracks the fix. One 40-day span punishes the team for finding a bug.
- A "Done" that means "ready for QA" measures half the path. Either call it development time or move the finish line.
- A blocked item should usually stay inside the interval, because the block is the delay you most need to see.
When you change the policy, do it on a dated cutover and keep the old series. Don't splice the charts together, since a new definition is not a delivery improvement. Pick the finish by who is waiting: if marketing is waiting on a launch, staging is not finished, and if support is waiting on a fix, production is the finish. Try the policy by hand on ten recent tickets. If people argue, it's usually because the column names are fuzzy, not because the math is wrong.
The formula
The formula for cycle time on a single item is:
Finish timestamp minus start timestamp.
Use one unit. Hours suit small items, and days are easier to say out loud. Don't mix "2.5 days" and "53 hours" in the same review. Use days with one decimal, or whole hours.
For the finished items in a window (a week, a sprint, or the last 30 items), report these:
- Quote the median. Sort the values and take the middle one. With an even count, average the two middle values.
- Add a high percentile, usually the 85th: walk 85 percent of the way up the sorted list. That gives you the "most items finish by" figure to use for a promise.
- Show the count next to the figure. A median of 20 items means more than a median of four.
- Include the average only if you have to. One ugly item pulls it up, so keep it out of the headline.
Don't average the percentiles of two teams and call the result a department number. Pool the raw items first, then compute.
A worked example
Eight items finished in two weeks. The policy is calendar time from "In progress" to "Released." The table shows hours; the published report uses days.
| Item | Started | Released | Elapsed |
|---|---|---|---|
| Tax rounding on invoices | Tue 10:00 | Thu 15:00 | 53 hours |
| Password reset copy | Tue 11:00 | Tue 16:00 | 5 hours |
| Export to CSV | Wed 09:00 | next Wed 09:00 | 168 hours |
| Empty-state illustration | Thu 14:00 | Fri 11:00 | 21 hours |
| Webhook retry | Fri 10:00 | Mon 10:00 | 72 hours |
| Billing-page crash | Mon 09:30 | Mon 13:30 | 4 hours |
| Audit log filters | Mon 15:00 | Thu 15:00 | 72 hours |
| Seat-count API | Tue 09:00 | the following Tue 09:00 | 168 hours |
Sorted hours: 4, 5, 21, 53, 72, 72, 168, 168.
The median is the average of the 4th and 5th values: (53 + 72) / 2 = 62.5 hours, about 2.6 days.
The 85th percentile sits near the 7th value (0.85 × 8 = 6.8): 168 hours, or 7 days. A promise of 2.6 days would be wrong about a quarter of the time. A promise of 7 days matches this sample.
The average is 70.4 hours, about 2.9 days, which is close to the median here. Add one 30-day item, though, and the average jumps while the median barely moves. That's why the median stays the headline.
The timestamps can't tell you whether the two seven-day items were large or just waiting, but the columns can. If "In review" held them for five of those days, the fix is review capacity. Adding another developer won't shorten a review-and-staging queue; two protected review slots a day will. The blended figure tells you a wait exists, and time in each column tells you where.
Benchmarks that hold up on your own board
There is no universal good number. A one-line studio change and a bank ledger migration shouldn't be measured against the same yardstick. A claim like "elite teams are under X hours" is usually a neighboring metric borrowed for the occasion.
The DORA research program groups lead time for changes, deployment frequency, change fail rate, and recovery time. DORA's lead time runs from commit to production. A board clock often starts earlier, at the first active status, so you can look strong on the DORA clock and still take a month to get started. Use DORA for release engineering and finished tickets for flow.
Build your own baseline from the last 20 to 50 finished items, by type:
- Compute the median and 85th percentile for incidents, small changes, and ordinary features.
- Label the result "us, this quarter." Compare next quarter against that, not against a blog post.
- If the median got shorter because you turned down large work, that's a portfolio change, so say so.
- Rebuild the baseline whenever the start column, stop column, or team boundary changes.
Illustrative ranges, not industry data
The ranges below are teaching examples. They are not survey results, and they are not targets.
- "One screen, one service" work might sit near a two-day median and a one-week 85th percentile. The conversation worth having is about the tail.
- Cross-service work with an external security review might sit near a two-week median. A two-day target won't speed up the review. It will push work off the board.
- Already-diagnosed defects might sit under a day, with a tail of vendor waits. Manage the tail.
If the figure is worse than last quarter, ask which column grew. If it's better, ask whether slow work dropped out of the sample. A target like "five days for everything" gets gamed with trivial tickets and by moving "done" to "ready for QA." Prefer a sentence you can check, such as "85 percent of accepted production defects are released within four days," measured the same way all year.
What moves the number
Elapsed time grows for a short list of reasons, and "we need more people" is usually fifth on it.
Too much work in progress: eight engineer-days spread across eight tickets finish later than the same days spent closing two. A WIP limit is the bluntest fix. How to set WIP limits, and what changes when you do covers the practical side.
Hidden queues: analysis, build, review, test, and deploy are often five separate waits all labeled "In progress." Split columns until each wait has a name.
Batch size: a weekly release train adds up to seven days to anything that missed it. A deliberate train belongs in the baseline. An accidental one is the biggest lever on the board.
Handoffs: a security review can earn its place, but many approvals are just habit. Ask which one changed a decision last month.
Item size: bigger work should take longer. Tag each item small, medium, or large when it starts, or you'll end up scolding a team for a quarter spent on bigger bets.
The constraint: if review is where the queue is, a developer who doesn't review only adds to the pile. Put the next hour of capacity on the column holding the most age.
Each week, record the median, the 85th percentile, the count, and the column with the longest wait. Then change that column, or write down that you're leaving it as is.
Reading cycle time next to the metrics people confuse with it
The number misleads when it's the only score you look at. Three neighboring metrics change how you read it.
Lead time starts when the request arrives and ends at the same finish. A short work clock with a long lead time means the wait is in the backlog, not in the build. Read The difference between lead time and this clock before you try to "fix" a team that is already fast once work starts.
Throughput is items finished per week. You can shorten elapsed time on the same number of items, or finish more items at the same elapsed time. Reading throughput and elapsed time together stops the trick of looking faster by doing less.
Takt time is the pace of demand: how often a new request arrives. If requests arrive faster than you finish them, elapsed time grows no matter how hard people work. A plain-English comparison of takt and this clock shows that mismatch.
Forecasts turn finish-time history into a range of dates. The 85th percentile is a crude forecast. A Monte Carlo simulation redraws that history thousands of times and returns a date with a probability attached. Monte Carlo forecasting for software delivery dates shows the mechanics, and Why delivery estimates miss, and how history fixes them covers the estimating habit those forecasts replace. When the business asks for a date, use one of those two.
If the date feeds a launch, give marketing the forecast. Launch dates marketing can plan against only exist when the history is stable enough to publish.
Flow charts that explain a move in the number
When the median jumps, the distribution shows the new shape, and a cumulative flow diagram shows which column stacked up. How to read a cumulative flow diagram pairs well with this metric: widening bands are queues, and a flattening finish line means throughput is dropping.
DORA metrics for a small team sit alongside this metric. Deployment frequency and change fail rate tell you whether a shorter clock made releases riskier. A faster board that ships more broken releases hasn't improved, so review them in the same meeting.
How the number gets distorted
One way is to start the clock later. Moving the start from "In progress" to "In review" deletes the development wait. Label review time as review time and keep the old series.
Another is hiding the slow kind of work. Vendor waits can be their own segment, but dropping them from the only chart leadership sees is how a dependency lingers for a year.
A hero week distorts it too. Six tiny tickets closed on Friday pull the average down, while the 85th percentile, and the feature a customer is waiting on, stay where they were.
Cancelled work is easy to miss. Abandoned items never finish, so they never enter the sample. Read the oldest open ticket alongside the median of finished work.
Then there are bonuses. Once pay depends on a shorter clock, the start column drifts and real work leaves the board. Review the figure, but don't pay people on it.
How to start this week
A spreadsheet is enough. So is Jira, Linear, Azure Boards, Shortcut, or any tracker that stores status changes. The policy matters more than the tool.
- Write down the start status, the stop status, which types count, and whether weekends count.
- Take the last 30 items that reached the stop. Record the start, stop, and type of each.
- Compute the median and the 85th percentile, overall and by type, as in the worked example.
- For the five longest items, name the slowest status. The answer has to be a column, not "people."
- For two weeks, change one thing: a WIP limit, a daily review slot, or a smaller batch. Then recompute.
Falrow's insights keep the percentiles, the cumulative flow, and a Monte Carlo forecast on the same ticket history. Still, if you can't explain the median on eight items by hand, don't trust the chart. A team that can't say where the clock starts doesn't have a metric yet.
FAQ
Does the clock run on weekends?
Yes, if you chose calendar time, which is the right default when a customer is waiting. A Friday-afternoon start that ships Monday morning is most of three days from the customer's side. Use business hours only for a staffing question, and label that chart so nobody compares it with the calendar series. State the choice in the policy and don't switch it from week to week.
Should you use the median or the average?
Use the median as the headline and a high percentile, often the 85th, as the promise. The average is optional. One month-long item drags it up and makes a normal week look like a collapse, or does the reverse once that item ages out of the window. In the worked example the median was 2.6 days and the 85th percentile was 7. An average near 2.9 days would have hidden the promise you can actually keep.
Can one slow item dominate the result?
It dominates the average. Once you have more than a handful of finishes, it barely moves the median. You should still name the slow item in the weekly review, find the column it sat in, and decide whether that wait is acceptable. Percentiles keep the outlier visible without letting it stand in for the typical item.
How many finished items do you need before the number means anything?
A dozen items of the same type gets you past pure noise, and around 20 to 50 makes a comfortable baseline. Below that, one unusual ticket can be the whole story. If you finish three items a week, use the last 20 finishes rather than a calendar week. Print the sample size next to the median so a quiet week can't pass for a trend.