Skip to content

Copilot Usage Metrics Dashboard Guide for Teams

Copilot Usage Metrics Dashboard shows adoption depth and delivery trends. Learn what to trust and how to act.

The Vibe Father 15 min read

The Copilot Usage Metrics Dashboard now tries to answer a better question than who opened the tool. GitHub's new impact view groups developers by how deeply they use Copilot and places delivery signals beside those groups. Enterprise administrators and organization owners can see passive users, code first users, agent first users, and people using multi agent or Copilot app workflows.

The dashboard, announced July 22, 2026, shows average pull requests merged per user, median merge speed, average lines of code per day, cohort size, six month trends, and an adoption multiplier that compares passive users with engaged users. It also recommends next steps aimed at moving users toward deeper product use.

That is useful information. It is not a scientific answer to whether Copilot caused a team to move faster. The view groups people by observed behavior and compares their outcomes. Experience, project type, team process, telemetry gaps, and selection effects can influence the same numbers. This guide explains what the dashboard measures, what it cannot prove, and how to use it without turning developers into chart targets.

The four adoption groups
Passive
Licensed and not engaged
Phase 1
Code first use
Phase 2
Agent first use
Phase 3
Multi agent or app use

Assignments use product activity from a rolling 28 day window

Who can see the new view

GitHub says the impact dashboard is available to enterprise administrators and organization owners who already have access to Copilot usage metrics. The wider metrics documentation also lists organization administrators, billing managers, and custom enterprise roles with the view permission for parts of the reporting system.

Access is not only a product setting. Usage data can reveal working patterns, tool preferences, and differences between teams. Decide who needs the dashboard and why. A small enablement group may need cohort trends. A line manager may not need individual level exports.

Document the purpose before opening the reports. A good purpose is improving training, licenses, and workflow support. A bad purpose is ranking individual developers from lines of code. The same data can support either behavior, so governance must come first.

How the cohorts work

GitHub uses its existing ai_adoption_phase field to assign groups from Copilot product activity across a rolling 28 day window. Passive users have a license but are not engaged. Phase 1 reflects code first behavior. Phase 2 reflects agent first behavior. Phase 3 includes multi agent or Copilot app behavior.

A cohort is a description of recent product use, not a skill level. Phase 3 does not mean a developer is better than someone in Phase 1. A database specialist, security reviewer, or maintainer of a sensitive system may have good reasons to use fewer agent features.

The 28 day window also means movement can lag real changes. A workshop held yesterday will not immediately produce a stable phase shift. A vacation, project transition, or incident week can change observed behavior without changing a person's attitude toward the tool.

What each cohort card shows

GitHub lists five main fields on the cards. Administrators see average pull requests merged per user per month, median pull request merge velocity, number of users, share of overall users, and average lines of code per day per user.

Together, these fields connect adoption depth with delivery activity. That is more informative than a seat count. It can show that many paid users remain inactive or that agent use is concentrated in a small part of the organization.

The fields also have different quality. User count is a direct adoption measure. Pull request speed is a workflow outcome shaped by review queues and change size. Lines of code is an activity measure that can rise when work becomes worse. Read them together and never let one field carry the whole conclusion.

A practical reading of the impact cards
MetricUseful questionMain caution
Users in phaseWhere is adoption concentratedPhase is not proficiency
Share of usersIs use broad or narrowLicenses may include poor fit roles
Pull requests mergedHow much work reaches mergeChange size and team role differ
Median merge velocityHow quickly review completesQueues and policy shape speed
Lines per dayHow much code activity appearsMore code is not more value

The adoption multiplier is not a causal estimate

The new dashboard compares the passive group with the average of engaged Copilot users and summarizes differences in throughput and speed. GitHub calls this the adoption multiplier. It gives leaders a quick view of how outcomes differ across the observed groups.

Do not translate that difference into Copilot made developers this much faster. People who choose advanced agent features may already be more experienced, work in repositories with better tests, receive more training, or handle different kinds of tasks. Passive users may include managers, occasional contributors, or people whose tools do not report rich telemetry.

Use the multiplier as a question generator. If Phase 2 merges faster, investigate why. Look at repository mix, pull request size, review staffing, tenure, language, and work type. Ask developers what changed. The comparison becomes valuable when it guides a careful follow up.

Why GitHub uses medians for merge speed

Pull request time can have huge outliers. A forgotten draft may stay open for months. A hotfix may merge in minutes. The median reduces the influence of unusually long requests and gives a more stable center than a simple average.

The median still hides distribution. Two teams can have the same median while one has predictable review and the other has many instant merges plus many stuck changes. Pair the dashboard with percentiles or a histogram from your engineering analytics when the decision is important.

Also check whether the metric starts at pull request creation. Developers may do more work before opening a request when an agent helps them. A faster merge phase does not necessarily mean the full idea to production cycle is faster.

Lines of code needs a warning label

GitHub describes lines of code measures as directional. They count code added or deleted through completions, chat actions, and agent edits. That can help reveal which modes produce activity. It should not become an output quota.

The best change may delete 2,000 lines. A weak generated solution may add 800 lines where 40 would do. Generated tests and vendored files can inflate totals. Languages also differ in verbosity. Comparing individuals by line count invites gaming and punishes maintenance work.

Use the measure to understand workflow shape. If agent initiated edits rise while review time and defects remain steady, the team may be absorbing the tool well. If lines soar while rollback and rework rise, the apparent productivity is probably debt.

Telemetry coverage shapes the story

GitHub says most usage details come from client telemetry in supported editors. Server side signals supplement that data and can count active users that client events miss. Network settings, proxies, user choices, and old extensions may reduce the detailed fields available.

Some surfaces are excluded from parts of the reports. GitHub documentation says the dashboard does not include every form of Copilot activity, and specific views may leave out command line use or activity on GitHub mobile and web chat. Read the definition for the chart you are using.

An active user detected only by a server signal may appear in top level counts while feature and line details remain empty. That person is not necessarily inactive or unproductive. They may simply have incomplete telemetry.

Data can arrive late

GitHub says usage reports update on a schedule and may take two full UTC days, while the dashboard viewing guide warns that some data may appear up to three full UTC days behind. Do not use the view like a live operations board.

Choose a consistent monthly reporting date that allows recent data to settle. Label the reporting window clearly. If a workshop ends on Friday, wait before judging its effect on Monday.

Late data also means a user can move phases after the period you are discussing. Save the export or snapshot that supports a decision. Otherwise a later visit may show a different classification and confuse the review.

Dashboard and interface reports can differ

GitHub says the dashboard, application interfaces, and export files use the same underlying telemetry but aggregate it differently. Enterprise counts may deduplicate a person while organization views can include the same licensed user in every organization they belong to.

Repository, organization, and user reports also have different shapes. Team level metrics are not prebuilt. GitHub instructs customers to join a user team report with the per user usage report when they need that view.

Small differences do not always signal a bug. Check scope, date window, time zone, granularity, late processing, and unknown values before opening a support ticket. Write those details beside every internal chart.

A responsible monthly review

A useful review can fit into one hour. Start with license activation and broad engagement. Then look at movement between phases. Examine delivery trends by team, but stop before drawing conclusions from a single month.

  1. Confirm data freshness and reporting scope
  2. Review active and passive seat trends
  3. Check phase movement over three months
  4. Compare delivery signals with repository context
  5. Read survey and retrospective feedback
  6. Select one enablement experiment
  7. Write the expected outcome and review date

Examples of a good experiment include a short agent mode workshop for one team, better repository instructions, a test suite improvement, or office hours for people who tried the tool and stopped. Change one thing where possible so the result is easier to interpret.

Report the learning, not only the number. If passive seats fell because managers were removed from the license pool, say that. If Phase 2 rose after a workshop but review time did not improve, say that too.

Pair the dashboard with developer feedback

Usage data tells you what buttons were used. It cannot tell you whether the developer trusted the result, felt interrupted, learned a new pattern, or spent an hour repairing generated code. A short recurring survey fills that gap.

Ask which task became easier, which task became harder, where the model wasted time, and what guardrail would make deeper use comfortable. Ask maintainers whether review quality changed. Keep responses anonymous when the topic could affect performance perceptions.

Compare sentiment with behavior. High activity and low trust may signal pressure to use the tool. Low activity and high interest may signal setup friction. Both patterns need a different response than generic training.

Do not turn phases into employee grades

A label such as Passive or Phase 1 can sound like a ranking. It is not. GitHub defines phases from product behavior, not impact, judgment, code quality, collaboration, or role fit.

Using the dashboard for individual performance reviews will distort the data. People will create unnecessary chat requests, accept suggestions they do not need, and choose agent mode to move a category. The metric stops describing work as soon as it becomes a target.

Keep evaluation at a group level unless a person asks for help with their own usage. Apply minimum group sizes in exported reports. Involve privacy, human resources, legal, and developer representatives before using user level data beyond support and license administration.

How to measure business value better

Start with outcomes the product team already values. Time from issue ready to production, escaped defects, incident recovery, review wait, customer support volume, and developer retention can matter more than raw code activity.

Use a balanced set. Adoption shows whether the tool reached people. Quality shows whether the work held up. Speed shows whether flow improved. Cost shows whether the gain was worth the spend. Experience shows whether the practice can last.

When possible, compare the same team before and after an enablement change and use a similar team as a reference. It still will not be a perfect experiment, but it is stronger than comparing advanced users with passive seats at one moment.

Use the interface for deeper analysis

The usage metrics interface and newline delimited export can provide more detail than the dashboard. Teams can examine models, languages, editor modes, repositories, and user level activity where policy permits. GitHub also documents fields for pull request lifecycle and code review adoption.

Build a small governed data model instead of copying exports into personal spreadsheets. Keep source date, scope, and field definitions. Limit retention. Remove direct identifiers when the question only needs a team trend.

Version your analysis. GitHub continues to add fields and surfaces, which can change totals or fill gaps. A chart should name the extraction date and documentation version so next quarter's analyst can reproduce it.

What leaders should put on one page

Show licensed seats, active share, phase distribution, accepted work trend, quality trend, direct cost, and one paragraph of developer feedback. Add the biggest data limitation in plain language.

Do not lead with the adoption multiplier alone. Put it beside the repository and role mix. Do not celebrate more lines without a quality measure. Do not hide a falling active rate behind a growing seat count.

End with one decision. Expand a workshop, retire unused seats, improve instructions, approve a model, or run a controlled trial. A dashboard earns its keep when it changes a useful decision.

The right way to use the new dashboard

GitHub's impact view is a welcome move beyond a simple active user counter. It helps teams see whether Copilot use is shallow or broad and connects those patterns to delivery. The six month trend can reveal whether enablement sticks after launch excitement fades.

The danger is treating a descriptive comparison as proof of causation or a product phase as a measure of talent. Resist both. Use the view to find where people need support, then confirm the story with workflow data and conversation.

A healthy program does not chase the highest possible phase. It helps each team use the right features for its work, with quality and privacy intact. The dashboard can guide that program when the organization keeps judgment in the room.

Sources and further reading

Common questions

What is the Copilot adoption multiplier

It compares pull request throughput and speed for passive users with the average of engaged users. It describes a difference between groups and does not prove Copilot caused that difference.

How often do adoption phases update

Assignments use product activity over a rolling 28 day window. Source data can also arrive several days after the activity.

Can managers use the dashboard for performance reviews

That is a poor use. The phases measure product behavior and miss role, code quality, project difficulty, collaboration, and telemetry gaps.

What should a team do with passive users

First ask why they are passive. The answer may be missing training, setup trouble, a role that does not benefit, or a license that should be reassigned.

The app behind this research

TheVibeFather is the multi-CLI AI coding harness

You just read field notes from the same team that ships TheVibeFather — the multi-CLI AI coding harness that runs Claude Code, Codex, OpenCode and more with shared memory and a verify gate. Bring your own keys.

Keep reading