> For the complete documentation index, see [llms.txt](https://wong-coupon.gitbook.io/flutter/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://wong-coupon.gitbook.io/flutter/my-flutter/quality-delivery/gitlab-ci-flutter-monorepo.md).

# GitLab CI for a Flutter Monorepo

Separate GitLab CI, Makefile, and Melos responsibilities to validate a Flutter monorepo consistently while preserving exit codes and reports on failure

## Result

In my Flutter monorepo, the pipeline does more than run `flutter test`. Before testing, the project must use the correct Flutter SDK, resolve dependencies at the workspace root, generate code in package order, and run the analyzer. CI must then retain coverage, logs, and test reports so I can identify the failed step.

After tracing `.gitlab-ci.yml` through the `Makefile`, Melos, and the reporting scripts, I separated the public workflow in this article into three layers:

```
                 GitLab pipeline
                       │
          trigger / runner / needs / artifact
                       │
          ┌────────────┴────────────┐
          │                         │
      Validate                    Test
      analyze           root + package + widget
          │                         │
          └────────────┬────────────┘
                       │
             report + raw artifact
                       │
                 quality result

Local and CI call the same Make targets
                       │
        Dart Workspace + Melos package graph
```

GitLab decides which jobs run, which runners execute them, and which artifacts they retain. Make provides commands shared by local development and CI. Dart Workspace and Melos manage dependencies and execution order across packages.

The observable outcomes are:

* Local development and CI start from the same Flutter version and Make targets.
* An analyzer failure or a failure in any test suite returns a non-zero job status.
* Raw logs and coverage remain available when the job fails.
* Once the project produces JUnit XML, GitLab displays the report in the pipeline instead of storing XML only as a downloadable file.
* Build and release jobs are added only after the quality checks pass.

This article does not claim that the current production pipeline is green. I verified the source, the Make call graph, and the configuration syntax; the GitLab Runner runtime, duration, and artifacts from a real pipeline still need verification on your GitLab instance.

## Problem

The original pipeline grew with the project. A real test job currently follows this sequence:

```
restore test dependency
→ make clean
→ bootstrap workspace
→ generate code
→ install Git hook
→ analyze
→ widget test
→ root/package test
→ merge coverage
→ upload coverage
```

The detail that caught my attention is that `make clean` does more than clean. It calls `make init`; `init` resolves dependencies, generates code, installs Git hooks, and writes Git metadata. The target name no longer reveals its actual side effects.

That creates four problems:

1. Developers cannot easily reproduce one failed step because a single command performs too many operations.
2. CI performs setup intended for developer machines, such as installing Git hooks.
3. A persistent runner can leave state in Git configuration or a shared dependency cache.
4. Every cache optimization or job split first requires untangling hidden side effects in the Makefile.

Test reports also have two different result layers. Uploading XML through `artifacts:paths` lets me download the file, but GitLab parses and displays a test summary only when the file is declared through `artifacts:reports:junit`. Even with JUnit, the exit code of the test command must still determine whether the job passes or fails.

The source also uses both `rules` and `only/except` in different jobs. GitLab keeps `only/except` for backward compatibility, but has deprecated them and recommends `rules`. With branches, tags, schedules, and manual pipelines, mixing both mechanisms makes job inclusion harder to verify.

The source contains BDD shard definitions, but their parent template currently sets `rules: when: never`; the Sonar job is manual and allowed to fail. I therefore do not describe BDD as actively sharded in CI or Sonar as a mandatory quality gate.

This article stops at the validate/test/report pipeline for Android and iOS. Fastlane, code signing, TestFlight, the Sonar quality gate, and go-live require separate articles because they use different runners, secrets, and publishing permissions from merge request validation.

## Solution

### Separate responsibilities before writing YAML

I use the following boundaries:

| Layer          | Owns                                                                  | Should not own                                            |
| -------------- | --------------------------------------------------------------------- | --------------------------------------------------------- |
| GitLab CI      | Triggers, runner capabilities, stages, `needs`, caches, and artifacts | Package loops or long code-generation logic               |
| Makefile       | Stable commands, call order, and exit codes                           | Internal branches, credentials, or production permissions |
| Dart Workspace | Workspace membership and dependency resolution                        | Pipeline inclusion conditions                             |
| Melos          | Scripts, filters, concurrency, and package order                      | Secrets or runner selection                               |
| Fastlane       | Build, signing, and upload after the quality gate                     | Baseline monorepo analysis and unit tests                 |

With these boundaries, `.gitlab-ci.yml` only calls `make ci-analyze` or `make ci-test`. Flutter logic is not duplicated in YAML, and the Makefile does not need to know whether the pipeline came from a merge request or a particular branch.

### Prepare the SDK, Workspace, and runner

The project I inspected uses Flutter `3.41.2`, Dart from `3.11.0` up to but not including `4.0.0`, Dart Workspace, and Melos `^7.4.0`. The Flutter version is pinned in `.fvmrc`:

```json
{
  "flutter": "3.41.2"
}
```

CI only needs to resolve the version already stored in the repository:

```bash
fvm install
fvm flutter --version
```

I do not hard-code the Flutter version again in `.gitlab-ci.yml`. When the SDK is upgraded, one change to the FVM configuration applies to both local development and CI.

If your project does not have FVM yet, the following article owns installation and Flutter version pinning. This article uses that result as a prerequisite.

{% content-ref url="/pages/757ESPKGR9SwyJPqwpIM" %}
[FVM](/flutter/my-flutter/foundations/fvm.md)
{% endcontent-ref %}

With Dart Workspace, one `dart pub get` at the root is the default contract. I do not repeat `pub get` in every package without a verified compatibility reason.

The following article covers Workspace declarations, `resolution: workspace`, and Melos scripts:

{% content-ref url="/pages/xhCc6xkN0EOAzEhnIdww" %}
[Dart Workspace and Melos](/flutter/my-flutter/architecture-state/dart-workspace-melos.md)
{% endcontent-ref %}

The runner needs the capabilities required by its job:

* The FVM binary to resolve the Flutter version pinned in the repository.
* Pure Dart/Flutter analysis and tests can run on Linux or macOS if the project dependencies support that platform.
* Android builds require a compatible Android SDK and JDK.
* iOS builds require macOS, Xcode, and CocoaPods.
* Reporting scripts may require Bash, Python, and LCOV.

The public example uses a neutral capability tag such as `flutter`. Real machine names, runner addresses, and infrastructure inventory do not belong in a public repository.

If a job does not use a container image, reproducibility also depends on how the runner is provisioned. FVM pins Flutter and Dart; it does not pin Xcode, Java, Ruby, Python, or LCOV.

### Split Make targets by responsibility

In the target configuration for this article, I split the target with excessive side effects into CI-specific targets. The reference source has not yet been refactored to match the following example:

```makefile
SHELL := /bin/bash

FLUTTER ?= fvm flutter
DART ?= fvm dart
MELOS ?= $(DART) run melos

.PHONY: ci-sdk ci-clean ci-bootstrap ci-generate ci-analyze ci-test ci-coverage

ci-sdk:
	fvm install
	$(FLUTTER) --version

ci-clean: ci-sdk
	$(FLUTTER) clean
	rm -rf build coverage reports

ci-bootstrap: ci-sdk
	$(DART) pub get

ci-generate: ci-bootstrap
	$(MELOS) run generate

ci-analyze: ci-generate
	$(FLUTTER) analyze

ci-test: ci-generate
	bash tool/ci/run_tests.sh

ci-coverage:
	bash tool/ci/merge_coverage.sh
```

Each target name describes its responsibility:

* `ci-clean` removes only reproducible project output.
* `ci-bootstrap` resolves dependencies from the workspace root.
* `ci-generate` runs the Melos script in dependency order.
* `ci-analyze` and `ci-test` are the entry points called by GitLab.
* `ci-coverage` only combines reports produced by tests.

These example targets do not install Git hooks, modify global Git configuration, or patch the dependency cache. Git hooks remain useful for developers, but CI does not need to install a hook before running a command that the pipeline calls directly.

The Make article explains target creation and local Make usage. Here, Make only acts as the command contract between the project and CI.

{% content-ref url="/pages/PG14RfyD6nuNYL68vG7H" %}
[Make](/flutter/my-flutter/foundations/make.md)
{% endcontent-ref %}

### Preserve exit codes while testing multiple packages

In a monorepo, I do not only check whether a `test/` directory exists. A package can contain placeholders only and make `flutter test` return `No tests were found`. The following script selects packages that contain at least one `*_test.dart` file:

```bash
#!/usr/bin/env bash
set -uo pipefail

PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
cd "$PROJECT_ROOT"

mkdir -p reports
rm -f reports/test-output.log reports/test-results-*.json

status=0

run_suite() {
  local name="$1"
  local directory="$2"

  (
    cd "$directory"
    fvm flutter test \
      --coverage \
      --reporter compact \
      --file-reporter "json:$PROJECT_ROOT/reports/test-results-$name.json"
  ) 2>&1 | tee -a "$PROJECT_ROOT/reports/test-output.log"

  local suite_status=${PIPESTATUS[0]}
  if [[ $suite_status -ne 0 ]]; then
    status=$suite_status
  fi
}

run_suite root "$PROJECT_ROOT"

for test_directory in packages/*/test; do
  [[ -d "$test_directory" ]] || continue
  find "$test_directory" -name '*_test.dart' -print -quit | grep -q . || continue

  package_directory="${test_directory%/test}"
  package_name="$(basename "$package_directory")"
  run_suite "$package_name" "$PROJECT_ROOT/$package_directory"
done

exit "$status"
```

Three details matter:

1. `set -o pipefail` and `PIPESTATUS[0]` preserve the status from `flutter test` instead of accidentally using the exit code from `tee`.
2. The script continues through the remaining packages to collect complete logs, but returns non-zero if any suite fails.
3. Old reports are removed before the new run so stale results cannot be mixed in.

If the project has a separate widget or BDD suite, I create another Make target instead of adding more conditions to the unit-test loop. The BDD article owns feature, step, World, and reporter organization; this article only owns the contract between the test command, its exit code, and GitLab artifacts.

{% content-ref url="/pages/5gKOCHnCinMIWgGm3uxB" %}
[BDD with Flutter Gherkin](/flutter/my-flutter/quality-delivery/bdd-flutter-gherkin.md)
{% endcontent-ref %}

### Build the baseline pipeline with `rules`

I start with merge requests and the default branch. When another branch pipeline is required, I add an intentional rule instead of hard-coding an internal convention into the article:

```yaml
workflow:
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
    - when: never

stages:
  - validate
  - test

default:
  interruptible: true

.flutter_job:
  tags:
    - flutter
  before_script:
    - fvm --version

analyze:
  extends: .flutter_job
  stage: validate
  script:
    - make ci-analyze

test:
  extends: .flutter_job
  stage: test
  script:
    - |
      status=0
      make ci-test || status=$?
      make ci-coverage || status=$?
      exit "$status"
  artifacts:
    when: always
    expire_in: 1 week
    paths:
      - reports/
      - coverage/
```

`workflow:rules` controls whether the entire pipeline is created, so jobs do not repeat the same conditions. `$CI_DEFAULT_BRANCH` keeps the example valid regardless of the repository's default branch name.

`test` follows stage order and begins only after validation. If analysis and testing are completely independent, you can place them in the same stage or use `needs: []` to start tests immediately. The trade-off is that each job must prepare dependencies and generated code independently or receive preparation artifacts through a clear contract.

I do not add `allow_failure: true` to mandatory quality checks. If analysis or tests are informational, the job name and policy should say so explicitly instead of leaving a green pipeline that readers might mistake for a passed gate.

### Declare JUnit for its actual role

Raw artifacts and structured reports do not replace each other. I retain both after the test infrastructure produces valid XML:

```yaml
test:
  artifacts:
    when: always
    expire_in: 1 week
    paths:
      - reports/
      - coverage/
    reports:
      junit: reports/**/*.xml
```

`artifacts:paths` retains logs, JSON, screenshots, and XML for download. `artifacts:reports:junit` tells GitLab to parse the XML and display it in the test summary and Tests tab.

JUnit does not fail the job. The test script must still return non-zero. Conversely, do not declare a JUnit path before a command actually creates the XML; the pipeline will only show a missing-artifact warning and provide no report for diagnosis.

I keep LCOV as the raw coverage artifact. For line annotations in a merge request, convert the report to Cobertura or JaCoCo and declare `artifacts:reports:coverage_report`. The `coverage` keyword only extracts a percentage from the log:

```yaml
test:
  coverage: '/\s*lines\.*:\s*([\d.]+%)/'
```

The coverage merge script must delete the previous aggregate file before appending package reports. Otherwise, a persistent workspace can carry coverage from an earlier run into the new pipeline.

### Cache only reproducible data

GitLab cache is appropriate for dependencies reused between jobs or pipelines. Artifacts are appropriate for output from a specific job. I do not use a cache as evidence that tests ran.

After the dependency cache is no longer patched in place, I can move the Pub cache inside the project and build its key from the SDK and lockfile:

```yaml
variables:
  PUB_CACHE: "$CI_PROJECT_DIR/.pub-cache"

.flutter_job:
  cache:
    key:
      files:
        - .fvmrc
        - pubspec.lock
    paths:
      - .pub-cache/
```

When `.fvmrc` or `pubspec.lock` changes, the cache key changes. I do not cache generated source without a mechanism that proves the output still matches its inputs.

The source I inspected still patches a package in the Pub cache to keep legacy tests running. In that state, sharing `.pub-cache` across jobs or branches can propagate modified state. I prefer removing the patch, pinning a reviewed fork or commit, or isolating the cache by version before enabling a shared cache.

Caching can reduce runtime, but I do not promise a number without measurements. I record separate durations for SDK setup, `pub get`, code generation, analysis, tests, coverage merging, and artifact upload before deciding to shard or increase concurrency.

### Keep secrets out of validation pipelines

The example pipeline needs no secrets. When a later step requires a coverage service, signing, or a store API, I keep the value outside the repository.

Sensitive GitLab CI/CD variables should be masked, hidden, and protected where appropriate. A protected variable only limits where GitLab provides it; it does not make merge request code trustworthy. Reviewers still need to inspect `.gitlab-ci.yml`, the Makefile, and scripts before allowing a job with secrets to run.

I avoid the following commands and patterns in jobs that have access to secrets:

* `env`, `export`, or debug tracing that prints the complete environment.
* Writing credentials to raw logs or artifacts.
* Downloading and immediately executing a `latest` script from the network.
* Giving merge request validation access to signing or release credentials.

The current source downloads and executes a coverage uploader directly from the network. I do not include that command in the public example. When using a third-party uploader, I pin its version and verify its checksum or signature according to the vendor documentation. If the requirement is only to retain evidence, I prefer native GitLab artifacts and reports.

Build, signing, and release jobs must remain separate from merge request validation and receive secrets only on protected refs or environments. Signing configuration and production permission setup are outside this article.

### Verify the pipeline

Before merging the YAML, I verify it in this order:

1. Run GitLab CI Lint on the target instance.
2. Run `fvm flutter --version` and confirm the project SDK.
3. Run `make ci-analyze` locally.
4. Run `make ci-test` and confirm that the root and packages with real tests are included.
5. Run `make ci-coverage` from a workspace without old reports.
6. Deliberately fail one test: the command must return non-zero while logs and reports remain available.
7. Delete the JUnit XML: the reporter or report-validation step must fail clearly.
8. Change `.fvmrc` or `pubspec.lock`: the cache key must be recreated.
9. Open an untrusted merge request: its jobs must not receive release secrets.

The minimum acceptance matrix is:

| Case                        | Job result | Expected artifact                                     |
| --------------------------- | ---------- | ----------------------------------------------------- |
| Analysis passes, tests pass | Green      | Logs, coverage, and JUnit if configured               |
| Analysis fails              | Red        | Analyzer log; tests and release do not run            |
| Root test fails             | Red        | Root JSON/log and any report already produced         |
| One package test fails      | Red        | Logs from executed suites without masking the failure |
| Report generator fails      | Red        | Raw test log for diagnosis                            |

I checked the source with a local YAML parser, shell syntax checks, and `make -n` for the call graph. These checks do not replace CI Lint or a real pipeline on the target runner.

### Roll back pipeline changes

During implementation, I change the pipeline in small steps: split the Make targets, migrate one job to a new target, add reports, and only then add caching. This order makes it easier to identify the change that broke the pipeline.

If splitting analysis and tests into multiple jobs causes runner failures or duplicates too much setup, I return to one `quality` job while keeping the Make targets separate:

```yaml
quality:
  extends: .flutter_job
  stage: test
  script:
    - make ci-analyze
    - make ci-test
```

Rolling back YAML does not mean removing exit-code or artifact contracts. Both must remain even when the pipeline temporarily returns to one sequential job.

I do not roll back by changing the quality job to `allow_failure: true` or adding `when: never` indefinitely. Those changes only hide failures and make pipeline status stop reflecting source quality.

### Common failures and trade-offs

#### The pipeline cannot be created because of `needs`

**Symptom:** GitLab reports that a job needs another job that is not present in the pipeline.

**Cause:** `rules` excluded the target job while `needs` still requires it.

**Fix:** align the rules for both jobs or use `optional: true` when the dependency is genuinely optional.

#### The job is green even though tests failed

**Symptom:** the log contains a failure, but the job reports success.

**Cause:** the pipeline uses the exit code from `tee`, a reporter, or a command that ran after the test.

**Fix:** enable `pipefail`, read `PIPESTATUS[0]`, and return non-zero when any suite or reporting step fails.

#### GitLab stores XML but displays no tests

**Symptom:** the XML can be downloaded from the artifacts, but the Tests tab is empty.

**Cause:** the XML is present only in `artifacts:paths` or does not follow JUnit format.

**Fix:** validate the XML and declare it through `artifacts:reports:junit`.

#### Coverage changes unexpectedly between pipelines

**Symptom:** files without tests still appear, or line counts are duplicated.

**Cause:** the aggregate LCOV file was appended to output from an earlier workspace run.

**Fix:** delete the aggregate report before merging and use only reports created by the current job.

#### The command works locally but fails on the runner

**Symptom:** the same Make target cannot find a command in CI or uses a different tool version.

**Cause:** FVM pins only Flutter; the runner has not pinned Java, Xcode, Ruby, Python, or LCOV.

**Fix:** manage the runner image or provisioning as code and print required tool versions at the start of the job without printing secrets.

#### The cache makes tests unstable

**Symptom:** a job fails only after a cache hit or when two branches use the same runner.

**Cause:** the cache contains mutable output, a patched dependency, or a key that does not include the SDK and lockfile.

**Fix:** disable the cache to verify the diagnosis, narrow its paths, change the key, and remove mutation from shared cached data.

As a trade-off, one sequential job is easier to debug but provides slower feedback. Parallel jobs reduce wall-clock time but repeat setup, consume more runner minutes, and complicate reporting. I shard tests only after measuring suite duration and defining boundaries that do not share devices, ports, or mutable caches.

### Verified versions

* Flutter: `3.41.2`.
* Dart: from `3.11.0` up to but not including `4.0.0`.
* Melos constraint: `^7.4.0`.
* Platform: Android and iOS; validation logic is shared, while build runners differ by platform.
* GitLab/GitLab Runner: the version of the actual instance has not been confirmed.
* Source and documentation verification date: 2026-09-02.

### References

* [GitLab — CI/CD YAML syntax reference](https://docs.gitlab.com/ci/yaml/)
* [GitLab — Make jobs start earlier with needs](https://docs.gitlab.com/ci/yaml/needs/)
* [GitLab — CI/CD artifacts reports types](https://docs.gitlab.com/ci/yaml/artifacts_reports/)
* [GitLab — Caching in GitLab CI/CD](https://docs.gitlab.com/ci/caching/)
* [GitLab — Deprecated keywords](https://docs.gitlab.com/ci/yaml/deprecated_keywords/)
* [GitLab — CI/CD variables](https://docs.gitlab.com/ci/variables/)
* [Dart — Pub workspaces](https://dart.dev/tools/pub/workspaces)
* [Melos — Workspace scripts](https://melos.invertase.dev/configuration/scripts)
* [Melos — Exec](https://melos.invertase.dev/commands/exec)
* [Codecov — Uploading reports and integrity checking](https://docs.codecov.com/docs/codecov-uploader)

## Conclusion

The most important part of GitLab CI in a Flutter monorepo is not the number of stages. A pipeline becomes trustworthy when each layer has one responsibility: GitLab orchestrates jobs and artifacts, Make preserves the command contract, and Dart Workspace with Melos manages the package graph.

I recommend keeping validation and testing separate from signing and release. This boundary allows a failure to return the correct status while retaining enough reports for diagnosis. If your project has only one package and one short test command, you do not need this entire structure yet; add each layer when the workflow becomes repetitive or difficult to reproduce.

[Buy Me a Coffee](https://buymeacoffee.com/ducmng12g) | [Support Me on Ko-fi](https://ko-fi.com/I2I81AEJG8)
