Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Performance

Measure each independently loadable leaf lint against compiler workloads. Category and aggregate libraries register their constituents and are excluded from the per-lint inventory. The harness verifies discovery against Cargo metadata and uses a shared workload generator.

Benchmarking

cargo xtask bench
cargo xtask bench --full

The default command performs a bounded inventory sweep: it samples every leaf lint with a prebuilt driver and small workloads using shortened warm-up and measurement windows. It retains Criterion’s normal sample collection and plots. Fast and full file-size results use separate Criterion groups, so their sampling modes are not compared as code changes. Use --full for longer statistical measurements before comparing a performance change. Reports include compiler startup and workload checking; they do not isolate only callback time.

Use the same pinned compiler, build profile, machine and workload sizes when comparing results. Keep external workload and toolchain changes separate from the measured implementation. Open target/criterion/report/index.html; the site publishes this index under benches/report/. Keep the generated directory structure so the index can resolve its chart and result links.

Build timings are omitted. Reliable comparisons need explicitly controlled warm and cold caches, dependency state, profile and host; this report measures compiler analysis with a prebuilt driver. Timings from an uncontrolled runner would mostly describe cache state.

Build costs

The Dylint libraries use compiler-private dependencies. Link them through dylint-link. Cargo’s build and target directories must agree because UI tests load the linked dynamic libraries from the target directory. The Nix shell clears an inherited shared build directory so each Dylint child can select its own target directory. A forced shared build directory places linker-named libraries outside Dylint’s lookup directory.

Development and CI builds disable incremental compilation and debug information to limit artifacts. For source-level backtraces, opt in with CARGO_PROFILE_DEV_DEBUG=line-tables-only; use CARGO_PROFILE_DEV_DEBUG=2 for full debugger information. These overrides rebuild affected artifacts, so use them for focused lint development. Existing artifacts from previous profiles remain on disk until explicitly cleaned.

Lint crates retain both cdylib and rlib outputs: Dylint loads the shared library, while category and aggregate crates link their constituents through Rust libraries. The rlib feature controls Dylint registration symbols, not Cargo’s output types. Helper crates already produce only rlib. Removing either output from a constituent breaks standalone loading or group composition.

Nix scopes the lint dependency cache to the aggregate lint library instead of the entire workspace. Workspace verification still covers all of its original targets. The catalog is generated through the existing xtask command.

Incremental optimization measurements

Local x86_64 Linux measurements on the 0.2.0 workspace layout with nightly-2026-07-15 (before the 0.3 layout refactor):

WorkloadBeforeAfter
Fresh cargo build -p perf --lib target directory336.2 MiB145.3 MiB
Loadable libraries in that build51.4 MiB5.9 MiB
bevy-asset-source-after-asset-plugin test executable71.1 MiB34.0 MiB
Compiler check of 500 locally allowed Markdown doc blocks492 ms55 ms

The build comparison changes only development debug information from line-tables-only to 0, using separate empty target/build directories and already downloaded dependencies. Allocated disk usage counts each inode once. One build of each profile took 17.0 and 15.7 seconds respectively; these are observations, not statistically established build-time improvements. Test executable sizes compare the same package and dependencies under both profiles.

The Markdown comparison uses identical no-debug profiles with and without the per-item lint-level guard. Each of 500 public functions has a locally allowed doc comment containing spaced emphasis. The lint remains enabled at crate level, and rumdl.toml enables MD037. The figures are medians of five alternating driver runs after warmup, including compiler startup and metadata emission. This measures skipped work on suppressed documentation, not a speedup for enabled Markdown checks or the entire lint suite. The UI regression checks local re-enabling, local suppression, and fulfilled expectations.

These measurements exclude toolchain installation, Nix store closures, and downstream project dependencies. Existing caches are not automatically removed.

Parser dependencies in focused builds

Workspace lints disable dylint-support’s default features. Only large-rust-file and unnecessary-module-directory enable its rust-file-size feature, which supplies the shared source-size parser. Direct consumers of the support crate retain that feature by default. This keeps unrelated lint builds from compiling the support crate’s full syn parser and span-location support. Other dependencies can still require syn, including compiler-side proc macros.

On the 0.3.2 workspace with the pinned compiler and debug information disabled, a fresh cargo build --locked --offline -p perf --lib used 145.3 MiB before this feature split and 126.4 MiB afterward, a 13% reduction. Both measurements used empty temporary target/build directories, warm downloaded dependencies, and one Cargo build job; allocated disk usage counted each inode once. The single builds took 29.1 and 25.2 seconds, respectively, which is insufficient to establish a repeatable build-time improvement. Temporary artifacts were removed.

These savings apply to focused builds that do not enable the parser elsewhere. The full aggregate and workspace test builds still include the parser-dependent lints, so they are not expected to see the same reduction.

Fixture dependencies

Bevy 0.19 fixture dependencies use umbrella re-exports and only required features. Bevy 0.18 remains for deliberate cross-version diagnostics. Bevy derive macros require the current umbrella dependency’s canonical name bevy; aliases are not discovered by its macro manifest helper.

Dylint’s upstream CI caches Cargo tools, registry and Git sources, compiler toolchains and Dylint drivers. The workflow also caches workspace build artifacts. Compiler and tool caches use their own versions rather than the workspace lockfile, so workspace release bumps do not invalidate them. Restore the compiler before invoking Rustup, and skip tool installation on an exact tool-cache hit. Compiler and tool caches are saved immediately after successful setup, even if a later gate fails. Registry and Git-source caching belongs to the workspace Cargo cache to avoid uploading the same data twice. Failed main-branch validations also retain dependency artifacts, so fixing a late gate does not force another cold dependency build. Keep coverage instrumentation outputs separate from ordinary build artifacts. After uploading coverage reports, CI discards the instrumented build, raw profiles and retained objects; the next measurement deliberately rebuilds those, so caching them would waste transfer time and disk space.

The Cargo CI job uses two compiler jobs on the public Ubuntu runner (four CPUs and 16 GB RAM), leaving headroom for compiler and native-dependency memory. Tests remain serialized, and local and Nix build limits remain unchanged. See GitHub runner specifications.

CI runs ordinary tests once with nextest, with doctests and instrumented coverage as separate gates. The small xtask integration suite also runs before workspace Clippy and UI tests so broken development commands fail early. Nix verification retains its independent sandboxed test run. The previous failing run spent 3½ minutes reinstalling tools despite restoring a cache; cache-hit savings and total CI times must be measured on successful runs of the updated workflow.

Source: Dylint 6.0.3 CI and Bevy 0.19 macro manifest lookup.

Attribute-free source-size checks

Source-size checks skip test-region parsing when a file contains no # character: every supported test-only marker requires an attribute. Sources containing # still use the full parser, including comments and strings that merely resemble attributes. Both the shared file-size helper and crate-size lint use this filter.

A development-profile microbenchmark of the shared file-size helper on 1,600 const _: () = (); lines took 22.65 ms before and 0.16 ms afterward per call. These are medians of seven batches of 30 calls using black_box, on the pinned compiler with debug information disabled. This isolates the helper and does not measure compiler startup, file I/O, or full-suite runtime.

Repeated-cfg diagnostic locations

repeated-cfg-gate indexes source line starts once per file that contains cfg gates. Each diagnostic endpoint then uses an indexed lookup instead of scanning all preceding lines. UTF-8 boundary checks, empty trailing lines, and invalid location rejection are preserved. The index uses one usize per source line.

A development-profile microbenchmark resolving column zero on each of 2,000 lines took 65.49 ms before and 0.18 ms afterward, including index construction. The figures are medians of seven batches, with inputs and results passed through black_box. This measures location conversion alone, not syntax parsing or a complete lint run; files with few gates have less opportunity to benefit.

Clap source snapshots

Clap derive and documentation recovery now borrows rustc’s retained source text instead of reopening and reading the complete file for each lookup. The shared source-map handle is cloned without cloning its text, and only the selected source segment is copied. This also keeps byte offsets aligned with the compiler’s source snapshot. Files without retained text still use the existing disk fallback.

Regression tests verify that retained text, including an empty snapshot, works without an accessible file, and that fallback reads and missing-file handling remain available. This removes file I/O and full-file allocations on the retained source path; no whole-suite runtime percentage has been measured for this change.

Candidate-only public-type scanning

unnecessary-public-type skips workspace scanning when the crate has no candidate types. Otherwise it stores only candidate-name counters, capped at two occurrences, and stops reading additional files when all candidates have a second occurrence. Identifier slices are borrowed instead of allocated per token. The existing lexical matching policy is unchanged.

A development-profile microbenchmark scanning 50,000 unrelated unique identifiers and two occurrences of one candidate retained 1 counter instead of 50,001. Median scanning time over seven batches was 52.9 ms before and 20.1 ms afterward. This excludes workspace discovery and file I/O; memory growth now follows the number of candidate names instead of all distinct workspace identifiers.

Shared cfg source text

repeated-cfg-gate now retains shared handles to rustc’s source strings instead of copying all loaded source text into a second collection. Files without # are omitted before parsing because they cannot contain cfg attributes. Files with attribute-like text still use the existing syntax-aware parser. The change removes full-source copies and parser work for attribute-free files; it does not change cfg classification or diagnostic locations. No whole-suite speedup is claimed for this change.

Coverage source identity caching

The coverage-based complexity lint resolves each reported source path once per compilation instead of canonicalizing it for every function. Successful and failed resolutions are cached in the report’s state; caches do not persist across compilations. Coverage still comes only from the measured executable-line records.

A development-profile microbenchmark of 5,000 coverage queries for one source file had median batch times of 49.1 ms before and 2.1 ms after caching, over seven batches. The report was parsed before timing. This measures repeated source lookup and fraction calculation, not report loading, complexity analysis, or compiler startup. It does not imply the same benefit for one-function-per-file workloads.

Bounded catalog discovery

Catalog discovery visits only crates/<crate>/ui and crates/<crate>/src, rather than recursively traversing fixture and build-output subtrees first. Source files inside a selected crate’s src are still read recursively. On the measured checkout, one discovery walk visited 1,647 entries instead of 3,169. Seven-run median times were 9.27 ms and 8.21 ms; the small timing difference is not evidence of a substantial catalog-generation speedup. The structural benefit is bounded traversal even when nested fixture/build directories grow.

Reqwest fixture features

Reqwest fixtures explicitly enable TLS, blocking requests, cookies, and multipart instead of enabling all defaults. Charset decoding, HTTP/2, and system-proxy integration are not used by these compile-time fixtures. The Linux Reqwest subtree drops from 131 to 127 distinct package/version entries: encoding_rs, fnv, h2, and tokio-util leave that subtree. The lockfile also drops unused platform-specific proxy dependencies. Packages required elsewhere can remain in the workspace graph; these counts are not a measurement of total disk savings. TLS diagnostics still compile against the existing default TLS implementation.

Router-only Axum fixtures

Axum fixture dependencies disable default features because the fixtures exercise routing and middleware APIs rather than server startup, JSON/form/query extractors, or tracing integration. The Linux Axum dependency subtree drops from 64 to 35 distinct package/version entries under the measured workspace resolution. Other workspace packages may still need some of those dependencies. All Axum UI fixtures retain their expected diagnostics with the smaller feature configuration.

Sparse field-cohesion graphs

The field-cohesion lint now connects each field’s users to one representative method and indexes direct callees, rather than comparing all method pairs and materializing shared-field cliques. Connected components—and therefore the existing diagnostic thresholds and summaries—remain the same. Regression tests compare components against the original pairwise algorithm across 4,096 field and call combinations, including self-calls and unknown callees.

For 300 methods sharing four fields, a development-profile microbenchmark of graph construction and component traversal stored 598 adjacency entries instead of 89,700. Seven-run median times were 74.64 ms before and 1.67 ms afterward. This isolates graph work; it excludes rustc traversal and does not represent a whole-suite speedup. Sparse graphs avoid quadratic edge storage for shared fields.