Abstract

A reasonable estimate can still be followed by a very late finish. In "Finish dates don't exist. Odds do.", I examined how widely completion times varied when software work shared the same recorded estimate and followed the same route through a project's workflow. This article explains the data and methods behind that comparison.

I used historical records from TAWOS, a public dataset of Jira issues from open-source software projects, to study Rubin Observatory Data Management, MongoDB's Evergreen system, and Appcelerator Studio. For each project, I reconstructed the progression from recorded start to completion using status histories, separated delivered work from abandoned work, and built comparison groups around the project's own workflow. Completion times were measured and compared separately within each project.

The spread remained substantial. Among 880 one-point Stories following the observatory project's standard route, half finished within about three days, while roughly five in every hundred took more than four weeks. The two company-operated projects showed the same pattern. In every group of at least a hundred issues, across all three projects, the late finish took at least three and a half times as long as the typical one. For anyone using estimates to set deadlines, the gap matters. A typical duration tells you where the middle was. The later finishes show how much risk that single number leaves out.

The question this analysis addressed

A missed date often leads to a request for better planning. Development is asked to estimate the work again, add the estimates together and choose a new deadline, usually with some extra margin. Sometimes the deadline comes first, and the estimates decide how much work should fit before it. In both versions, the estimates carry the weight of the deadline.

I wanted to examine the relationship that calculation depends on: how much does a task's estimate tell us about its time to completion?

For the main comparison, I selected planned development work, leaving bugs out, and compared completed work with the same estimate within each project. I then narrowed the comparison to work that passed through the same stages in the same order.

I measured elapsed calendar time from the recorded start of development activities to completion. This includes waiting and review as well as development work. A deadline has to account for the whole interval.

The dataset and the records available

TAWOS and the version used here

I used TAWOS, a dataset of public Jira records from open-source software projects. Version 1.1 contains 458,232 issues from 39 projects, including recorded estimates and histories of status changes. Those histories let me reconstruct how work moved from start to completion. The records used here run up to 23 October 2020. [1, 6, 7]

TAWOS was assembled from projects with evidence of Agile practice and at least 200 issues with story-point estimates. The dataset paper describes version 1.0, which also included five Talendforge projects later removed when their tracker ceased to be public. Before the analysis, I audited the imported records against that paper. The data-quality findings from that audit are explained below. [1, 6]

Three projects, and why these three

I selected three projects with enough estimated, completed work to support substantial comparison groups and workflows I could interpret. Together, they cover scientific software and two tools run by software companies, with different ways of organizing the work. [6, 8-11]

  • Rubin Observatory Data Management (DM) builds the software that turns the observatory's sky images into science data. Its published workflow provided a clear starting point for interpreting the records. [9]
  • MongoDB Evergreen (EVG) is the system MongoDB uses to test software changes across platforms. [10]
  • Appcelerator Studio (TISTUD) was a tool for building mobile apps. [11]

I interpreted each project's records separately. The project-specific sections below explain which work entered the comparisons and which periods were used. [7, 8]

How the analysis is put together

I kept the imported records unchanged and stored my reconstructions separately. The analysis follows three steps:

  1. Configuration files define what each project's statuses and resolutions mean, and which work and routes enter the comparison.
  2. The builder turns the recorded status changes into a timeline and a route for each issue.
  3. The study script selects the comparison groups and calculates their completion-time summaries. [7, 8]

Jira supplies the status-change timestamps; my reconstruction rules determine the start and completion events. I set each project's status and route mappings before running the completion-time study. The following sections explain those rules and the corrections needed to make the records usable.

What the source papers established

The closest prior study compared developers' estimates, recorded as story points, with the time issues spent in progress. The relationship was usually weak or moderate, and issues with the same estimate showed substantial variation in duration. Some took much longer than others, even after the researchers removed very short durations and statistical outliers. [2]

Two other papers tested methods for predicting the story-point estimate an issue would receive. In both studies, a prediction using the middle value of earlier estimates was competitive with more elaborate methods. Predicting an assigned estimate and predicting a finish date answer different questions. [3-5]

Alongside the TAWOS dataset paper, these studies informed the data audit and the choice of time measure. [1-4] My comparison focuses on the full interval from development start to completion, including waiting and review, and groups work by both estimate and workflow route. It retains the long durations because those late finishes matter for deadlines. [7, 8]

Why the supplied duration columns were insufficient

TAWOS includes calculated duration fields, but I needed the elapsed time from development start to completion. The audit showed that the supplied fields did not consistently measure that interval. [6]

  • Resolution time often included the backlog. For 83 percent of resolved issues with a recorded move into In Progress, the value matched the time from ticket creation to resolution. That includes time before development starts, although the dataset paper describes the field as starting when work begins. [1, 6]
  • Time in progress depended on the workflow. The supplied value was zero for 76 percent of resolved issues. Almost all of those tickets still had status changes; the zero often reflected a mismatch between the field's status definition and the project's workflow. Requiring a positive value would have selected work according to the status names used. [6]

I therefore reconstructed start and completion from each project's status history, interpreting its stages separately. The next section explains those rules. [8]

What the recorded estimate could establish

The supplied estimation date repeated the ticket's creation time, sometimes one hour earlier. It could not establish whether an estimate had been made before development began. [6]

For the stable-estimate comparisons, I selected estimates from 1 to 100 and required the revision flag (Story_Point_Changed_After_Estimation) to be unset. Some issues had several estimate entries in their history despite an unset flag, so I describe this restriction as "no revision recorded by the flag." [1, 2, 6, 7]

Timestamp precision

The timestamps showed seasonal one-hour discrepancies consistent with local clock changes, despite the documentation describing them as UTC. I applied no timezone correction and report durations in elapsed calendar days. An interval crossing a daylight-saving change can therefore carry a one-hour discrepancy. This matters most for groups where half the work finished in less than a day, and for completion times close to the whole-day boundaries used in the charts. [1, 6, 8]

Reconstructing an issue's recorded lifecycle

I reconstructed each Jira issue's timeline from its dated status changes to identify the start, completion and route used in the comparison. [8]

Start and completion

  • Start: first entry into a status marking work as under way.
  • Completion: first entry into the final uninterrupted sequence of terminal states, which mark work as ended.

For example, an issue that moves from In Progress to Resolved and later to Closed finishes at the move to Resolved. The later closure adds no time. If it reopens before finishing again, I use that later completion and include all the time in between. [8]

I call the interval from recorded start to final completion cycle time. It includes development, review and waiting, plus time spent in earlier terminal states before a reopening. The interval from ticket creation to the first active state is the pre-start wait and sits outside cycle time.

For the study, I selected delivered issues with a recorded resolution date and a measurable start. Open work and work with no observed active state were outside the completion-time comparison. [7, 8]

Cycle time runs from the first entry into an active status to the first entry into the final run of terminal statuses. Time before a reopen counts; the pre-start wait and a later administrative closure do not. Status names from Appcelerator Studio; the same rules apply in all three projects.
Cycle time runs from the first entry into an active status to the first entry into the final run of terminal statuses. Time before a reopen counts; the pre-start wait and a later administrative closure do not. Status names from Appcelerator Studio; the same rules apply in all three projects.

Classifying the stages

Status names differ between projects, so I mapped them separately to six classes. Each continuous period in an active status is an active spell. [8]

  • active — work marked as under way, such as In Progress.
  • queue — waiting, including backlog states, blocks and waits after review.
  • review — undergoing review or waiting for a reviewer.
  • done — the status label marks the work as delivered.
  • abandoned — the status label marks the work as not delivered.
  • terminal — neutral closing states, such as Resolved or Closed; the recorded resolution determines delivery.

A status such as DM's Reviewed means review has finished, so I classed it as queue. An issue in the neutral terminal class is treated as open without a recorded resolution date. An unmapped resolution on a resolved issue in that class stops the build. [8]

Within a cycle, I summed time in active, review and queue states separately. Earlier time in terminal states before a reopen is stored as prior terminal dwell. These intervals add up to cycle time; the study reports the total elapsed interval. [8]

Main flow and standard route

A route is the sequence of status classes an issue passed through, with consecutive occurrences of the same class collapsed into one step. For example, queue → active → review → queue → done means waiting, work, review, another wait and delivery. [8]

  • Main flow: the routes I accepted for comparison because they ended in delivery and passed through the stages I treated as necessary in that workflow. Rework, requeue and reopen loops remain eligible.
  • Standard route: the most common eligible route in the project. Within that route, grouping by estimate holds both the estimate and stage sequence constant. [7, 8]

For DM, the handbook's statuses Todo, In Progress, In Review, Reviewed and Done map onto the standard route. The guide defines one point as an idealised half day of work. I use that current documentation as context; the historical workflow decisions come from the project investigation. [8, 9]

Handling history inconsistencies

I ordered each issue's status changes by timestamp, with changelog ID breaking ties. The first genuine status change's origin status applies from ticket creation. Each later interval uses the destination status of the change that opened it. Without genuine changes, I use the current recorded status from creation. [8]

Entries such as Closed → Closed are ignored because they record no change of state. If consecutive events disagree about a state, I retain the label that opened the interval, flag it as disputed and include the interval in the analysis. [8]

I revised the completion boundary and the rule for unchanged status labels after examining the outcomes, then applied both rules to all three projects. The Appcelerator Studio section explains what each correction changed. [8]

The project-specific decisions

The same comparison needed different selection rules in each project.

DM: Stories with recorded work and review

I used the full recorded period, from 2014 to October 2020. DM's Stories included code changes alongside operations, reporting and accounting work, so the issue type alone was not enough to select the process I wanted to compare. I required both an active spell and a recorded review before delivery. [7-9]

The review requirement excludes the large queue → active → done route, which included work reviewed on GitHub without entering Jira's review stage. It also excludes queue → active → queue → done, including issues that moved through In Progress → Reviewed → Done: Reviewed is grouped as queue. [8]

Epics and Milestones were excluded because their timelines aggregate other issues. I kept Technical tasks separate because missing parent links prevented a check for overlapping parent and child timelines. Bugs and other execution types appear in separate checks. [8]

I split the results at 1 October 2019, after a workflow change removed the direct move from review to Done. Era 0 is before that boundary; era 1 starts there. I report both the full period and the eras separately, assigning each issue by its reconstructed completion date. The route queue → active → review → done stays eligible in the main flow in both eras; the standard route includes a wait after review. [7, 8]

EVG: separating delivery from cancellation

I selected Task, Improvement and New Feature. I checked that Task was being used for standalone engineering before including it; the result files also show the comparison without Task. Bugs were examined separately. [8]

Both active work and review were required. I grouped In Progress, Debugging and Investigating as active, and In Code Review as review. Closed contained delivered, cancelled and duplicate issues, so the recorded resolution determines delivery. Deployed is grouped as queue because issues normally moved on from it to closure. The endpoint measures completion of the recorded workflow. [8]

The comparison uses era 2, from 5 February 2018 onward. This is the period after Closed replaced Resolved as the normal ending status. [7, 8]

TISTUD: separating completion from later closure

I selected Story, Improvement and New Feature over the full period, from March 2011 to October 2020. Review was optional in this workflow, so I required active work and delivery. Bugs and Technical tasks were examined separately. [7, 8]

The history contained 2,219 Closed → Closed entries and later administrative moves from Resolved to Closed. An earlier rule ended the cycle at the last terminal event. Under that rule, the time by which about 95 percent had finished was near two and a half years. Using the first entry into the final uninterrupted terminal run removed that administrative delay; genuine reopens still extend the measured lifecycle. [7, 8]

I also ignored repeated-label entries so they could not create extra intervals or route steps. I applied the rules for completion and repeated labels to all three projects. Ignoring repeated-label entries left DM and EVG's cohort results unchanged to two decimal places. [7, 8]

I retained the project's issue-type labels, even where a Story's title looked like a bug report. No era split was adopted: after repeated-label entries were ignored, the mapped workflow retained the same stage structure. [8]

From project records to comparison cohorts

Selecting the comparison groups

I built the comparison groups in three steps:

  1. Keep delivered issues with a recorded resolution date and a measurable cycle time of zero or more minutes. For EVG, use only era 2.
  2. Apply the project's selected issue types and main-flow routes.
  3. Keep estimates from 1 to 100, including fractions within that range, with no revision recorded by the flag. Then select the standard route and group by exact estimate. [6, 7]

The table starts with delivered issues across all types. DM's starting count already requires a measurable cycle; EVG's and TISTUD's do not. Every later row applies that requirement. [7, 8]

Selection stepDMEVG, from 2018TISTUD
Delivered issues before type and route selection14,1274,2023,433
Selected types on a main-flow route, with measurable cycle time5,9971,629906
Same, with estimates from 1 to 100 and the revision flag unset3,7781,428775
Same, on the standard route3,1441,030414

Table 1. Counts as each project's population was narrowed, using the issue types selected in the project-specific sections above. [7, 8]

The largest reduction came before filtering estimates, when I applied the requirements for issue types, routes and a measurable cycle. Those requirements excluded DM's skip-review route, EVG issues without a recorded work start or review, and TISTUD issues that closed without a recorded start. [8]

The standard route was different in each project:

  • DM: queue → active → review → queue → done
  • EVG: queue → active → review → terminal
  • TISTUD: queue → active → terminal

Within each standard route, I grouped issues by their exact estimate. DM's groups contain only Stories. EVG and TISTUD combine their selected issue types, so the estimate and route match within a group while the type can differ. [7]

How the published counts fit together

The route figure in the original article covers 14,210 DM Stories with estimates from 1 to 100, including work that was delivered, abandoned or still open at the cutoff. The completion-time comparison here uses 880 delivered one-point DM Stories with a measurable cycle, no revision recorded by the flag, and the standard route. The 880 are among the 14,210, but the counts came from separate queries for different comparisons. [7, 12]

DM's 14,210 estimated Stories followed 176 routes after status labels were grouped into stages, with 4,122 on the standard route. The original route figure reports 3,853 following the handbook's exact status sequence, but the query behind that stricter count was not retained. [12]

How to read the completion times

I used percentiles to compare typical and late finishes:

  • P50 (the median): the time by which about half the group had finished.
  • P85: the time by which about 85 percent had finished.
  • P95: the time by which about 95 percent had finished.

The ratio P95/P50 expresses the late completion point as a multiple of the median. A ratio of 3.5 means that point was three and a half times the median. [7]

I calculated percentiles from unrounded durations and report them in elapsed calendar days, including nights and weekends. Durations and ratios are shown to two decimals, with ratios calculated before rounding. Headline comparisons use groups of at least 100 issues. [7]

Results and checks

DM: variation remained within point groups

Among DM's 3,778 stable-estimate main-flow Stories, the median was 8.15 days and P95 was 88.63 days, a ratio of 10.87. Restricting to the standard route left 3,144 Stories, with a median of 7.17 days, a P95 of 71.34 days, and a ratio of 9.96. Holding the point value constant as well, the 880 one-point Stories have a ratio of 10.18. Matching the route and then the estimate moved the centre of the distribution but barely touched its shape. [7]

Completion times for 880 delivered one-point DM Stories on the standard route, with the revision flag unset, over the full study period. Finishes beyond day 30 are grouped together in the chart and remain in the percentile calculations. [12]
Completion times for 880 delivered one-point DM Stories on the standard route, with the revision flag unset, over the full study period. Finishes beyond day 30 are grouped together in the chart and remain in the percentile calculations. [12]

Holding the recorded point value constant produces the following groups with at least 100 observations.

DM recorded pointsIssuesP50, daysP85, daysP95, daysP95/P50
18802.7412.8727.8710.18
27496.0223.0546.887.79
32649.9833.9375.527.57
44519.0832.9060.036.61
513413.0142.14102.057.84
621815.2543.0181.015.31
814423.0454.2788.353.84
1013425.8772.10117.184.53

Table 2. DM standard-route point groups, pooled eras. All groups use Story, the stable-estimate restriction, delivered measurable cycles, and the standard route. Ratios use unrounded values. [7]

The central duration rises with the estimate, except that 4-point Stories had a slightly lower median than 3-point ones. The 95th percentile remains several times the median in every group. The shared estimate and route did not make observed durations tightly concentrated.

The time by which about 95 percent had finished, as a multiple of each group's median. Each block is one median: for one-point Stories a block is 2.7 days, and the 95 percent point lies just over ten blocks out. Delivered DM Stories on the standard route with the revision flag unset, full study period; every group has at least 100 issues.
The time by which about 95 percent had finished, as a multiple of each group's median. Each block is one median: for one-point Stories a block is 2.7 days, and the 95 percent point lies just over ten blocks out. Delivered DM Stories on the standard route with the revision flag unset, full study period; every group has at least 100 issues.

Workflow-era checks

The obvious objection to the pooled DM result is that it straddles a workflow migration. I therefore split the results at the October 2019 boundary. On the standard route, era 0 holds 1,981 Stories running from a median of 8.84 days to a P95 of 78.41 days, a ratio of 8.87; era 1 holds 1,163 running from 5.25 to 57.63 days, a ratio of 10.98. For the one-point standard-route cohort, era 0 holds 478 issues running from 3.07 to 37.91 days, a ratio of 12.34, and era 1 holds 402 running from 2.02 to 20.09 days, a ratio of 9.94. The spread appears in both eras, not only in the pooled figure. [7]

The era comparison does not show that the migration caused any change in completion times. Era assignment follows the endpoint, and other conditions could have changed. The check addresses a narrower concern: substantial spread remains when the eras are examined separately.

EVG and TISTUD under their own rules

The company-operated projects use their own type populations and standard routes. Their point scales are not a shared unit of work across projects.

Project and recorded pointsIssuesP50, daysP95, daysP95/P50
EVG, 13911.1813.9111.75
EVG, 24592.9715.115.08
EVG, 41455.1819.363.74
TISTUD, 51550.244.3517.84
TISTUD, 81261.767.564.31

Table 3. Standard-route point groups with at least 100 observations. EVG uses its study era and Task + Improvement + New Feature; TISTUD uses Story + Improvement + New Feature. Separate project results, not pooled estimates. [7]

EVG's 1,428 stable-estimate main-flow issues had a median of 4.77 days and a P95 of 48.99 days. One objection is that the Task type might be driving the result. The sensitivity check excluding Task kept 786 Improvement and New Feature issues, with a median of 5.27 days and a P95 of 68.01 days. The spread was present without Task. [7]

TISTUD's main flow contained 775 stable-estimate issues, with a median of 2.06 days and a P95 of 100.41 days. Its standard-route cohort contained 414 issues, with a median of 0.81 days and a P95 of 7.30 days. The two summaries differ so much largely because the main flow also admits the reopen routes that the standard route excludes: the 93 stable-estimated issues on queue → active → terminal → queue → terminal run from a median of 16.15 days to a P95 of 202.17 days. [7]

The five-point TISTUD cohort has a sub-day median. Its large ratio is sensitive to a small denominator, and timestamp precision matters more at that scale. The eight-point cohort gives the clearer day-scale comparison: a median of 1.76 days and a P95 of 7.56 days.

Across all reported comparisons with at least 100 issues, the lowest P95/P50 was 3.53, for DM's eight-point Stories in era 0. [7]

What the results mean for planning

Across three projects, matching the estimate and workflow route still left a substantial gap between typical and late finishes. Waiting, review and rework are part of the elapsed interval a deadline has to account for. [7, 8]

I compared completed issues using recorded estimates and the routes each issue eventually took. I did not test how reliably the historical percentiles predict future completion times. [7, 8]

For planning, a typical duration is a useful reference. Historical late finishes show the risk that reference leaves out. Asking how likely work is to finish by the date you need makes that risk part of the decision.

References

TAWOS is public. My project mappings, scripts and result files are unpublished working materials, described in references 6–8 and 12.

  1. Tawosi, V., Al-Subaihin, A., Moussa, R., and Sarro, F. (2022). A Versatile Dataset of Agile Open Source Software Projects. MSR 2022, doi:10.1145/3524842.3528029. Dataset provenance and field descriptions, especially §§2.1 and 2.5. The data and its README live in the TAWOS repository; this analysis uses the version 1.1 database download, doi:10.5522/04/21308124.
  2. Tawosi, V., Moussa, R., and Sarro, F. (2022). On the Relationship Between Story Points and Development Effort in Agile Open-Source Software. ESEM 2022. Time definitions in §2.2, sample and filters in §3.3, results in §4.
  3. Tawosi, V., Moussa, R., and Sarro, F. (2022). Agile Effort Estimation: Have We Solved the Problem Yet? Insights From A Replication Study. IEEE Transactions on Software Engineering accepted manuscript, arXiv version 2, 17 December 2022. Experimental results and replication discussion.
  4. Tawosi, V., Al-Subaihin, A., and Sarro, F. (2022). Investigating the Effectiveness of Clustering for Story Point Estimation. SANER 2022. Dataset and chronological split in §IV-B, comparisons in §V.
  5. Choetkiertikul, M., Dam, H. K., Tran, T., Pham, T., Ghose, A., and Menzies, T. (2019). A Deep Learning Model for Estimating Story Points. IEEE Transactions on Software Engineering 45(7). The estimator that reference 3 replicates.
  6. TAWOS import and data-quality investigation, 2026. Unpublished working notes by the author: version reconciliation against the dataset paper, the resolution-time and estimation-date column discrepancies, the timestamp offsets, missing fields, and known data limitations. Investigation records, not a peer-reviewed publication.
  7. Completion-time study, 2026. Unpublished analysis by the author: a method note, a manifest recording the 23 October 2020 cutoff, the study script, per-project result files for DM, EVG, and TISTUD, and findings notes for each project recording interpretation and earlier checks. Tables 1, 2, and 3 are taken from the result files.
  8. Workflow reconstruction, 2026. Unpublished analysis by the author: a method note, the builder, and for each of DM, EVG, and TISTUD the adopted status, type, resolution, route, and era configurations with the investigation and validation records behind them.
  9. LSST Data Management, the Rubin Observatory's software project under its former name, the Large Synoptic Survey Telescope. DM Development Workflow and Data Management Project Management Guide, DMTN-020, both accessed 8 September 2026. Current documentation supplies project context, including the five workflow steps and the idealised half-day story point; the historical workflow interpretations used here come from the investigations in reference 8.
  10. MongoDB (27 July 2016). Evergreen Continuous Integration: Why We Reinvented the Wheel, accessed 8 September 2026. Context for the EVG project.
  11. Axway. Changes to Application Development Services, accessed 8 September 2026. Context for Appcelerator Studio and its discontinuation after the dataset's cutoff.
  12. Figure materials for the original article, 2026. Unpublished, held with the analysis: the completion-day data and its extraction, the bar-figure script, the route figure, the estimate-range figure, and the full listing of the 176 routes taken by the 14,210 estimated Stories.