The Copilot Studio Kit: Test Agents Before Employees Do
How the Copilot Studio Kit builds test sets for your own agents, what test types exist, and how to install and configure it.
The Copilot Studio Kit, officially Copilot Agent Kit, is a complementary tool maintained by the Power Customer Advisory Team (Power CAT) for Microsoft Copilot Studio that lets you test, monitor, and secure your own agents before employees or customers see them. The core of the kit is test sets that confront an agent with defined inputs and validate responses automatically against expected values instead of manually clicking through each change in chat. As of August 2026.
How do you install the Copilot Studio Kit?
The easiest way is through the Marketplace: Visit the Copilot Agent Kit in Marketplace select Get it now, sign in, and select the target environment.
- Marketplace Installation: Confirm terms, select Install, then in Power Apps under Solutions open the default solution and configure the Copilot Studio Kit - Dataverse connection reference with a valid Dataverse connection.
- GitHub Installation: Download the latest managed solution from the official Power CAT repository and import it in Power Apps under Solutions, recreating connections as needed.
- Access: After installation, you'll find the app under Apps as Power CAT Copilot Studio Kit, open it by selecting Play, and then share it with your team using the Share function.
A setup wizard in the kit handles configuration of connection references, environment variables, and cloud flows mostly automatically.
How do you build test sets for your agents?
A test set bundles multiple individual tests that you run together against an agent. The core building blocks are six test types with different validation depths.
- Response Match: compares the agent response directly with an expected answer, optionally exact, contains, or starts with.
- Attachments: validates structured attachments like Adaptive Cards, optionally using AI-assisted validation against custom validation instructions.
- Topic Match: compares expected with actually triggered topics and supports multiple topics simultaneously in generative orchestration.
- Generative Answers: uses a language model to assess whether an AI-generated answer matches a reference answer or custom validation instructions.
- Multi-turn: chains multiple individual tests in fixed order within the same conversation and stops on critical failures.
- Plan Validation: checks in generative orchestration whether the agent actually uses expected tools in its execution plan, evaluated against a pass threshold in percent.
Test sets can be created and maintained in bulk through Excel, and individual tests or entire sets can be duplicated to quickly derive variations.
What does the kit do beyond pure testing?
Beyond test automation, the kit includes several modules that address governance and fleet operation of agents. The Compliance Hub evaluates the agent inventory against configured controls and flags policy violations, the agent debugger module provides step-by-step diagnostics for recorded conversations, and Conversation KPIs aggregate usage and quality metrics alongside the built-in Copilot Studio analytics. For organizations using Copilot Studio in production, test automation and Compliance Hub in combination are the most effective lever against uncontrolled agent proliferation.
What is the connection test useful for with protected data sources?
When an agent accesses an external data source like SharePoint or Dataverse, an authorization card appears in the first conversation that a user must confirm. You replicate this flow using a multi-turn test with two child tests: the first test triggers the authorization card and compares its content, the second sends confirmation and validates the subsequent response. Once authorized for a test user, later tests against the same connector work without this step. Those looking to embed this test coverage into an existing automation strategy with Microsoft Copilot benefit from a clear separation between development and production environments.
Frequently Asked Questions About Copilot Studio Kit
Does Copilot Studio Kit cost additional licensing fees?
The kit itself is a free complementary tool maintained by Microsoft's Power CAT team that you can install from the Marketplace or GitHub. However, it requires a Power Apps and Dataverse environment, which follow standard Power Platform licensing rules.
Can I run tests automatically before every deployment?
Yes, through integration with Power Platform pipelines, tests can be run automatically before an agent is promoted between environments. The kit can thus be used for multi-stage release processes from development through testing to production.
Which test type is suitable for a multi-step conversation flow?
The multi-turn test type models multiple successive messages in the same conversation and can be combined with Response Match, Attachment, Topic Match, and Generative Answers tests as child tests. Critical child tests stop the entire test run on failure, non-critical tests continue and provide additional context.
Do I need additional configuration for Generative Answers tests?
Yes, this test type requires AI Builder enrichment to be enabled so a language model can evaluate the generated answer against a reference answer or custom validation instructions. Without this enrichment, the test type is not available in the interface.
Simon Glowik
Founder of NordFlux. Spent four years automating processes at enterprise scale at Dräger, and now brings that depth to the mid-market — pragmatic and with full data sovereignty.
Certifications
- Microsoft certified — PL-900 and AZ-900
- UiPath certified — Automation Developer Associate
- UiPath zertifiziert — Automation Developer Associate
Concrete questions about automation or AI?
In a free initial analysis we discuss your case directly. No strings attached.