> For the complete documentation index, see [llms.txt](https://wong-coupon.gitbook.io/flutter/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://wong-coupon.gitbook.io/flutter/my-flutter/quality-delivery/startup-scrolling-performance-benchmark.md).

# Benchmarking Startup and Scrolling Performance

How I separate startup traces, scrolling timelines, and runtime telemetry, then turn artifacts into fail-closed quality gates on real Flutter devices

## Outcome

I do not use one number to represent “app performance.” An app launch, a scrolling action, and the experience on a user's device are three different measurements.

I split the system into three paths:

| Measurement path    | Question                                                                       | Output                        | Feedback mechanism   |
| ------------------- | ------------------------------------------------------------------------------ | ----------------------------- | -------------------- |
| Startup benchmark   | How long does first-frame build or rasterization take on a lab device?         | Startup JSON and raw timeline | CI exit code         |
| Scrolling benchmark | What build/raster distribution does a fixed scrolling action produce?          | Timeline and summary JSON     | CI exit code         |
| Runtime telemetry   | How do production traces change across devices, OS versions, and app versions? | Firebase custom trace         | Dashboard and alerts |

```
Scheduled CI on physical devices
    │
    ├── Startup benchmark
    │      ├── flutter run --trace-startup --profile
    │      ├── start_up_info_<platform>.json
    │      └── startup gate ──► exit code
    │
    └── Scrolling benchmark
           ├── deterministic integration scenario
           ├── traceAction
           ├── timeline_summary_<platform>.json
           └── scrolling gate ──► exit code

Production application
    └── Firebase custom trace ──► segmented trends
```

Both CI benchmarks run in profile mode on physical Android and iOS devices. Runtime traces collect data on user devices. These paths complement each other, but they cannot replace each other or be compared directly.

A principle that matters more than any threshold is **fail closed**:

```
artifact exists?
    └── valid JSON?
          └── required fields with correct types?
                └── sample count > 0?
                      └── metric within threshold?
                            ├── yes ──► pass
                            └── no  ──► fail
```

If a file is absent, JSON is missing a field, or a timeline contains no frames, the job must fail before threshold comparison. A script that skips a missing file and exits with `0` is not a quality gate.

## Problem

### “Startup time” has multiple milestones

When I run `flutter run --trace-startup`, Flutter writes startup milestones to `start_up_info.json`. With Flutter `3.41.2`, the important fields include:

```
engineEnterTimestampMicros
timeToFrameworkInitMicros
timeAfterFrameworkInitMicros
timeToFirstFrameMicros
timeToFirstFrameRasterizedMicros
```

The three related formulas are:

```
timeToFrameworkInitMicros
  = frameworkInitTimestamp - engineEnterTimestamp

timeAfterFrameworkInitMicros
  = firstFrameBuiltTimestamp - frameworkInitTimestamp

timeToFirstFrameMicros
  = firstFrameBuiltTimestamp - engineEnterTimestamp
```

Therefore:

```
timeToFrameworkInitMicros + timeAfterFrameworkInitMicros
  = timeToFirstFrameMicros
```

This sum is the time until the first frame is **built**. It is not `timeToFirstFrameRasterizedMicros`, and it certainly does not prove that a screen has finished loading data or is ready for every interaction.

I name metrics after the event that was actually recorded:

* `first_frame_built`: the framework built the first frame.
* `first_frame_rasterized`: the engine rasterized the first frame.
* `interactive`: used only when the app has a separate marker for an interactive-state contract.

If I do not have an interactive marker, I do not rename a first-frame metric to make it appear stronger.

### One launch is not a distribution

One startup trace gives me one observation from one device in its state at that moment. It does not automatically provide a median, p90, variance, or confidence interval.

The result can be affected by:

* Whether the app or process is cold or warm.
* App data and persistent cache.
* OS version, device model, and refresh rate.
* Thermal state, battery state, and background workload.
* Flutter version, app version, and native flavor.
* Whether the app was just installed or only relaunched.

I do not mix cold install, cold process, and warm process samples in one distribution. Each scenario has its own name, reset rule, and baseline.

### A scrolling timeline is meaningful only when the action is stable

A scrolling test can include fixture setup, network waits, animations, and cleanup. If the trace wraps all of them, the summary no longer answers only “is scrolling smooth?”

I prepare the widget, service mocks, and finders first. Only the measured action belongs inside `traceAction`:

```
fixture + mock + initial pump
            │
            └── outside the trace

traceAction
    ├── scroll to the last item
    ├── wait for animations to settle
    ├── scroll back to the beginning
    └── wait for animations to settle
```

Putting `pumpAndSettle()` inside the callback means the timeline includes every frame until the scheduler settles. That can be the intended contract, but it must stay consistent across runs.

### Build and raster are different phases

`TimelineSummary` separates framework and engine time:

* Frame build covers widget building, layout, paint, and compositing.
* Frame rasterizer converts a scene into pixels.

The metrics answer different questions:

| Metric               | Meaning                                 | Limitation                                   |
| -------------------- | --------------------------------------- | -------------------------------------------- |
| Average build/raster | Average cost                            | Can hide tail latency                        |
| Worst build/raster   | Slowest frame                           | Sensitive to one outlier                     |
| p90/p99              | Tail of the distribution                | Requires enough samples                      |
| Budget-miss count    | Number of durations over a fixed budget | Depends on frame count                       |
| Frame count          | Sample size                             | Does not say whether frames are fast or slow |

In the SDK version I inspected, `TimelineSummary` uses a fixed `16 ms` budget for missed-budget metrics and explicitly notes that this budget does not account for the device's actual refresh rate. Therefore, `missed_frame_*_budget_count` is not the number of frames that certainly dropped on every display.

I always read the count together with `frame_count`. Two traces with different lengths should not be compared only through a raw missed count.

### Missing artifacts can create false passes

A common anti-pattern is checking a file only when it exists:

```bash
if [[ -f "$summary" ]]; then
  check_metrics "$summary"
fi

echo "All metrics passed"
exit 0
```

If a previous command did not produce the summary, the block is skipped and the script still reports success. A startup parser can also turn empty command-substitution output into `0` if it does not stop when `cat` or `jq` fails.

I do not rely vaguely on `set -e`. The parser explicitly verifies that:

* Every required platform artifact was received.
* The file is non-empty and readable.
* JSON parsing succeeds.
* Required fields exist, are numbers, and are non-negative.
* Frame and rasterizer counts are greater than `0`.
* Thresholds exist and have the correct unit and type.

Every precondition failure must return a non-zero exit code.

### Artifacts matter only if they survive job failures

GitLab runs `after_script` before artifact upload. If a job declares paths under `build/` but `after_script` moves those files elsewhere, the runner might no longer find them at the declared paths.

Artifacts are also uploaded only on successful jobs by default. This is the opposite of what an investigation needs: when a threshold fails, I need the raw timeline most.

I keep one canonical output directory and use:

```yaml
artifacts:
  when: always
  paths:
    - build/performance/
```

I do not use `after_script` for invariants that affect pass or fail. Under GitLab's default behavior, an `after_script` error does not turn an already successful `script` into a failed job.

### A Firebase trace is not a scrolling timeline

A Firebase custom trace measures duration between `start()` and `stop()` and can carry custom metrics. If a page starts a trace in `initState` and stops it in a post-frame callback, that trace describes one interval in the page lifecycle.

It does not directly give me:

* Frame build time.
* Rasterizer time.
* Frame p90/p99 for a scrolling action.
* A deterministic assertion that returns a CI exit code.

A post-frame callback also does not prove that remote data has loaded or that the page is interactive. I use runtime traces to study production trends, not to replace integration benchmarks.

The typed trace boundary, per-operation session lifecycle, and distinction between source evidence and Firebase Console evidence are covered separately here:

{% content-ref url="/pages/6qARxZcQKpqR3JK7pxwm" %}
[Firebase Performance and Custom Traces](/flutter/my-flutter/security-observability/firebase-performance-custom-trace.md)
{% endcontent-ref %}

## Solution

### Normalize the environment before measuring

Flutter recommends profiling performance on physical Android and iOS devices. Debug mode has asserts, JIT, and overhead that differ from release behavior; simulators and emulators do not represent the same hardware either.

I pin these inputs in the metadata for every run:

```yaml
benchmark:
  app_commit: <commit-sha>
  flutter_version: 3.41.2
  platform: android
  os_version: <os-version>
  device_class: <device-class>
  refresh_rate_hz: <refresh-rate>
  flavor: <benchmark-flavor>
  scenario: cold_process
  attempt: <attempt-index>
```

A device serial does not belong in public logs or widely shared artifacts. CI receives it through protected runner configuration or variables.

Thresholds belong to a platform, device class, and scenario. I do not use one threshold for Android, iOS, 60 Hz devices, and 120 Hz devices.

### Measure startup with `--trace-startup`

The minimal command is:

```bash
flutter run \
  --trace-startup \
  --profile \
  --target lib/main.dart \
  --flavor "$BENCHMARK_FLAVOR" \
  --device-id "$PHYSICAL_DEVICE_ID"
```

Flutter writes output to `build/` by default. If `FLUTTER_TEST_OUTPUTS_DIR` is set, the tool writes there instead.

In addition to `start_up_info.json`, Flutter creates `start_up_timeline.json`. I preserve both for each platform:

```
build/performance/
├── android/
│   ├── start_up_info.json
│   └── start_up_timeline.json
└── ios/
    ├── start_up_info.json
    └── start_up_timeline.json
```

The summary makes gating fast. The raw timeline can be reopened in a trace viewer when a regression occurs.

### Write a fail-closed startup parser

The parser reads one consistently selected metric. The following example gates first-frame rasterization; if the current baseline uses first-frame build, I retain that metric until I establish a new baseline.

```bash
#!/usr/bin/env bash
set -euo pipefail

require_startup_metric() {
  local file="$1"
  local key="$2"

  [[ -s "$file" ]] || {
    echo "Missing or empty startup artifact: $file" >&2
    return 1
  }

  jq -er --arg key "$key" '
    .[$key]
    | select(type == "number" and . >= 0)
  ' "$file"
}

android_value="$(
  require_startup_metric \
    build/performance/android/start_up_info.json \
    timeToFirstFrameRasterizedMicros
)"

ios_value="$(
  require_startup_metric \
    build/performance/ios/start_up_info.json \
    timeToFirstFrameRasterizedMicros
)"
```

Only after both platforms are read do I compare thresholds:

```bash
status=0

if (( android_value > ANDROID_STARTUP_LIMIT_MICROS )); then
  echo "Android startup exceeded its limit" >&2
  status=1
fi

if (( ios_value > IOS_STARTUP_LIMIT_MICROS )); then
  echo "iOS startup exceeded its limit" >&2
  status=1
fi

exit "$status"
```

Thresholds come from reviewed configuration; the script does not hard-code device serials or secrets. A public diagnostic includes the platform, metric, value, threshold, and unit, but no device identifier.

If I read both built and rasterized milestones, I add this invariant:

```
timeToFirstFrameRasterizedMicros >= timeToFirstFrameMicros
```

An invalid schema must fail. It must not coerce `null`, strings, or absent fields to `0`.

### Separate cold, warm, and repeated runs

I define reset rules for each scenario:

| Scenario     | State before launch                                              | Question answered                  |
| ------------ | ---------------------------------------------------------------- | ---------------------------------- |
| Cold install | The app was just installed and has no local state                | First-run cost                     |
| Cold process | The process was terminated while app data is retained            | Typical launch after process death |
| Warm process | The process or app state is warm according to a defined contract | Resume or relaunch cost            |

I run `<repeat-count>` attempts for one scenario, retain every sample, and then compute the median and a tail percentile. The repeat count comes from observed variance and lab cost; there is no universal number for every team.

I do not “rerun until green.” Repetition quantifies noise. If the distribution shifts, a threshold changes only after review and a new baseline.

### Build a deterministic scrolling scenario

The test uses `IntegrationTestWidgetsFlutterBinding` and service mocks. Fixture setup finishes before tracing:

```dart
final binding = IntegrationTestWidgetsFlutterBinding.ensureInitialized();

testWidgets('scroll performance', (tester) async {
  await pumpDeterministicScreen(
    tester,
    itemCount: fixedItemCount,
  );

  final list = find.byKey(const Key('benchmark-list'));
  final lastItem = find.byKey(const Key('last-item'));
  final header = find.byKey(const Key('header'));

  expect(list, findsOneWidget);
  expect(lastItem, findsOneWidget);
  expect(header, findsOneWidget);

  await binding.traceAction(
    () async {
      await tester.scrollUntilVisible(
        lastItem,
        scrollStep,
        scrollable: list,
        duration: stepDuration,
      );
      await tester.pumpAndSettle();

      await tester.scrollUntilVisible(
        header,
        -scrollStep,
        scrollable: list,
        duration: stepDuration,
      );
      await tester.pumpAndSettle();
    },
    reportKey: 'scrolling_timeline',
  );
});
```

These inputs must be fixed:

* Item count and fixture content.
* Scroll step, direction, and duration.
* Theme, locale, viewport, and relevant native flavor.
* Service responses and image providers.
* Animations and background tasks included in or excluded from the trace.

A finder failure is a functional failure and must stop the test before threshold parsing. I do not turn “unable to scroll” into an empty timeline and call it good performance.

Widget and golden tests retain their own behavioral and visual contracts. A performance test does not replace either layer.

{% content-ref url="/pages/OXzEEJ7gphdUcYxkfusW" %}
[Widget Tests and Golden Regression](/flutter/my-flutter/quality-delivery/widget-test-golden-regression.md)
{% endcontent-ref %}

### Write the timeline and summary on the host

`traceAction` places the timeline in `reportData` under its `reportKey`. The host driver reads that key and writes files:

```dart
Future<void> main() {
  return integrationDriver(
    responseDataCallback: (data) async {
      if (data == null || data['scrolling_timeline'] == null) {
        throw StateError('Missing scrolling_timeline report data');
      }

      final timeline = driver.Timeline.fromJson(
        data['scrolling_timeline'] as Map<String, dynamic>,
      );
      final summary = driver.TimelineSummary.summarize(timeline);

      await summary.writeTimelineToFile(
        'scrolling_timeline',
        pretty: true,
        includeSummary: true,
      );
    },
  );
}
```

The default output is:

```
build/scrolling_timeline.timeline.json
build/scrolling_timeline.timeline_summary.json
```

I set a separate `FLUTTER_TEST_OUTPUTS_DIR` for each platform so raw timelines cannot overwrite each other:

```bash
FLUTTER_TEST_OUTPUTS_DIR=build/performance/android \
flutter drive \
  --driver test_driver/perf_driver.dart \
  --target performance_profiling/scrolling_perf_test.dart \
  --flavor "$BENCHMARK_FLAVOR" \
  --device-id "$ANDROID_DEVICE_ID" \
  --profile \
  --no-dds
```

iOS uses the same target, driver, scenario, and flavor contract; only the device ID and platform-specific build change.

`--no-dds` is part of the harness I verified. I recheck this flag when upgrading Flutter instead of treating it as a permanent contract.

### Choose scrolling metrics

My minimum gate validates the sample before any metric:

```
frame_count > 0
frame_rasterizer_count > 0
```

Only then does it read build and raster metrics. One policy might use:

* p90 build time to protect the tail of UI work.
* p90 rasterizer time to protect the tail of GPU/raster work.
* Worst time as a wider outlier guard.
* A missed-budget ratio instead of a raw count when trace length can vary.
* A frame-count range to detect major action-boundary changes.

An example schema check with `jq`:

```bash
jq -e '
  (.frame_count | type == "number" and . > 0) and
  (.frame_rasterizer_count | type == "number" and . > 0) and
  (."90th_percentile_frame_build_time_millis"
    | type == "number" and . >= 0) and
  (."90th_percentile_frame_rasterizer_time_millis"
    | type == "number" and . >= 0)
' "$summary" >/dev/null
```

After schema validation, threshold comparison can use `jq` itself instead of parsing floating-point numbers through shell integer arithmetic:

```bash
jq -e \
  --argjson buildLimit "$P90_BUILD_LIMIT_MS" \
  --argjson rasterLimit "$P90_RASTER_LIMIT_MS" '
    ."90th_percentile_frame_build_time_millis" <= $buildLimit and
    ."90th_percentile_frame_rasterizer_time_millis" <= $rasterLimit
  ' "$summary" >/dev/null
```

I do not gate every `TimelineSummary` field. A chosen metric needs an owner, a baseline, and a failure action. Adding more thresholds does not automatically make a benchmark trustworthy.

Flutter's `integration_test` also provides `watchPerformance`, described as a migration path away from `traceAction` and `flutter_driver.TimelineSummary`. I migrate only after running both paths together and establishing a parity baseline; changing the measurement API and thresholds at the same time destroys the meaning of the trend.

### Wire scheduled CI

Performance benchmarks are long-running and depend on physical devices, so I run them through a schedule or explicit `JOB_TYPE` instead of on every merge request:

```yaml
performance_scrolling:
  stage: test
  rules:
    - if: '$JOB_TYPE == "scheduled scrolling benchmark"'
  script:
    - make benchmark-scrolling
  artifacts:
    when: always
    expire_in: <retention-window>
    paths:
      - build/performance/android/
      - build/performance/ios/
```

Startup uses a separate job because its commands, metrics, and failure modes differ from scrolling.

If a runner has one shared device, I serialize jobs with an appropriate resource lock. A device occupied by another job is an infrastructure failure, not a performance regression.

I keep threshold parsing in `script`; `after_script` does not decide pass or fail. `after_script` is suitable only for cleanup or enrichment that cannot change the main conclusion.

The separate article about GitLab CI for a monorepo owns the broader discussion of caching and failure semantics:

{% content-ref url="/pages/iXaQgqQ1eqF71OtJywau" %}
[GitLab CI for a Flutter Monorepo](/flutter/my-flutter/quality-delivery/gitlab-ci-flutter-monorepo.md)
{% endcontent-ref %}

### Read failures at the correct layer

| Signal                                      | Common cause                                                            | Response                                                  |
| ------------------------------------------- | ----------------------------------------------------------------------- | --------------------------------------------------------- |
| Job was not created                         | Schedule or `JOB_TYPE` did not match                                    | Call it skipped or not scheduled, not passed              |
| Device not found                            | Lab configuration or offline device                                     | Fix infrastructure; do not change thresholds              |
| Startup missing engine/first-frame event    | Launch or VM service was interrupted                                    | Preserve the Flutter tool failure and raw log             |
| Startup JSON missing or empty               | Command did not generate or copy the artifact                           | Fail the precondition                                     |
| Summary JSON malformed                      | Driver or output is corrupt                                             | Fail schema validation and retain the raw timeline        |
| `frame_count == 0`                          | Action produced no frames or timeline scope is wrong                    | Fix the scenario or driver                                |
| Build percentile increased                  | Expensive UI-thread work, layout, or paint                              | Open the timeline and find the relevant frames and stacks |
| Raster percentile increased                 | Expensive scene, shader, image, or GPU work                             | Inspect raster and GPU events                             |
| Only one worst frame increased              | Outlier, GC, shader, or lab noise                                       | Inspect the raw timeline and repeated-run distribution    |
| Firebase trend increased while CI is stable | Production segment, device, network, or content is outside the scenario | Segment telemetry before reproducing it                   |
| Threshold failed without artifacts          | Artifact path or upload condition is wrong                              | Fix CI retention before investigating the app             |

### Keep a separate contract for runtime telemetry

A runtime flow can look like this:

```
State.initState
    │
    └── start custom trace
           ├── render/load work
           ├── add low-cardinality custom metric/attribute
           └── stop trace at a meaningful marker
```

If the marker is a post-frame callback, I name the trace after “first post-frame after page init,” not “page fully loaded.” If the product needs time to content or time to interactive, I add a marker at the correct state and test that marker's lifecycle.

Firebase custom traces already have a duration. A custom metric adds a value; it does not replace duration and must not contain PII. Trace names do not contain account IDs, route parameters, or dynamic identifiers.

### Benchmark review checklist

```
Environment
□ Physical device and profile mode
□ Flutter/app version, platform, OS, flavor, scenario
□ Device class/refresh rate and reset rule

Scenario
□ Deterministic fixture
□ Setup outside the trace
□ Fixed action boundary
□ Clear finder and functional-assertion failures

Artifacts
□ Raw trace + summary for each platform
□ No Android/iOS overwrite
□ Upload even when the job fails
□ No device serial or secret

Gate
□ File/schema/type/sample validation fails closed
□ Metric has a unit, baseline, and owner
□ Threshold changes are reviewed
□ Exit code reflects every platform

Conclusion
□ Do not extrapolate one scenario to the whole app
□ Do not call first frame “interactive”
□ Do not mix runtime telemetry with a CI benchmark
```

### Trade-offs and limitations

* Physical devices and profile mode are closer to production but make CI more expensive and harder to parallelize.
* Service mocks stabilize the action but do not measure production network or content.
* One screen benchmark does not represent every list in the app.
* Averages are easy to read but hide the tail; percentiles need enough samples.
* Worst values catch outliers but are sensitive to noise.
* `TimelineSummary`'s fixed `16 ms` budget does not account for the actual refresh rate.
* Repetition reduces the influence of one sample but does not repair a bad scenario.
* Thresholds catch large regressions but do not replace raw-timeline investigation.
* Firebase telemetry has production coverage but is not deterministic and does not return a CI exit code.
* `watchPerformance` is a migration path, not a reason to change a baseline silently.

### Verification performed

In the source snapshot I inspected:

* Flutter is pinned to `3.41.2`, with Dart `3.11.0`.
* Profile-mode startup commands exist for physical Android and iOS devices.
* Each platform currently has one startup attempt in the scheduled job.
* One mocked scrolling integration scenario scrolls down and back to the beginning.
* The host driver writes a timeline and `TimelineSummary`.
* The scrolling parser currently gates five metrics per platform.
* Both jobs are created only for matching `JOB_TYPE` values, not as default MR gates.
* Seven page states use the runtime performance mixin.
* `packages/performance` has no unit tests.
* The local snapshot has no benchmark artifact from which to publish a result.

I ran shell syntax checks, parsed the GitLab YAML, and dry-ran the Make commands. I also ran both parsers in a temporary directory without artifacts: the current startup and scrolling scripts both exited with `0`. This article therefore teaches fail-closed gating instead of claiming that the current pipeline is complete.

I did not run a benchmark on physical devices, validate thresholds against real artifacts, or inspect the Firebase console during research. No startup or FPS result was inferred from source code.

### Related article

A performance integration test uses the same driver foundation as functional E2E, but it has a different action boundary and different artifacts:

{% content-ref url="/pages/5gKOCHnCinMIWgGm3uxB" %}
[BDD with Flutter Gherkin](/flutter/my-flutter/quality-delivery/bdd-flutter-gherkin.md)
{% endcontent-ref %}

### Verified versions

* Flutter: `3.41.2`.
* Dart: `3.11.0`.
* `integration_test`: Flutter SDK.
* `flutter_driver.TimelineSummary`: Flutter SDK.
* `firebase_performance`: `0.11.2` in the lockfile.
* Platform scope: Android and iOS.

### References

* [Flutter performance profiling](https://docs.flutter.dev/perf/ui-performance)
* [Measure performance with an integration test](https://docs.flutter.dev/cookbook/testing/integration/profiling)
* [Flutter API — `traceAction`](https://api.flutter.dev/flutter/package-integration_test_integration_test/IntegrationTestWidgetsFlutterBinding/traceAction.html)
* [Flutter API — `TimelineSummary.summaryJson`](https://api.flutter.dev/flutter/flutter_driver/TimelineSummary/summaryJson.html)
* [Flutter API — `watchPerformance`](https://api.flutter.dev/flutter/package-integration_test_integration_test/IntegrationTestWidgetsFlutterBinding/watchPerformance.html)
* [Flutter startup tracing source](https://github.com/flutter/flutter/blob/90673a4eef/packages/flutter_tools/lib/src/tracing.dart)
* [Firebase — Custom code traces](https://firebase.google.com/docs/perf-mon/custom-code-traces?platform=flutter)
* [GitLab CI/CD — `after_script`](https://docs.gitlab.com/ci/yaml/#after_script)
* [GitLab — Job artifact upload conditions](https://docs.gitlab.com/ci/jobs/job_artifacts/#with-upload-conditions)

## Conclusion

Startup, scrolling, and runtime telemetry are three measurements with different boundaries, metrics, and failure semantics. I keep them separate so a regression can be localized to launch, UI build, rasterization, or a production segment instead of being compressed into the label “the app is slow.”

A trustworthy benchmark begins before the threshold: physical devices, profile mode, a fixed scenario, platform-specific artifacts, and a fail-closed parser. Missing files, invalid schemas, or empty samples must fail; they must not become `0` and pass.

The threshold is only the final step. Raw timelines, distributions, and metadata provide the evidence needed to explain why a metric changed. Firebase traces add the perspective of real users, but they do not replace deterministic CI benchmarks or exit codes.

[Buy Me a Coffee](https://buymeacoffee.com/ducmng12g) | [Support Me on Ko-fi](https://ko-fi.com/I2I81AEJG8)
