Trainable Classifiers for Records Management: A closer look

This article covers what trainable classifiers are, and where they can best serve your records management requirements.

Paula McClure

Paula McClure

|

Records Manager

Posted on: September 30, 2024

|

Last updated: June 25, 2025

|

10 min read

Paula McClure

Paula McClure

|

Records Manager

Posted on: September 30, 2024

|

Last updated: June 25, 2025

Read time: 10 minutes

Records Managers are always seeking ways to make filing easier for users. One way of achieving this is by utilising trainable classifiers in Microsoft 365, which use machine learning to automatically recognise and classify documents wherever they are stored.

In this blog, I’ll explain:

Ultimately, I hope you come away from this article with a greater understanding of what trainable classifiers are, and where they can best serve your records management requirements.

What is a Microsoft Purview trainable classifier?

A trainable classifier is a tool that can be used to identify and organise (classify) unstructured data, and can be trained with machine learning to learn from both positive and negatives matches given in sample content. Once trained, these classifiers can be used to apply Office sensitivity labels, communication compliance policies, and retention label policies to identified items.

Trainable classifiers screenshot 1

The uses and advantages of trainable classifiers

The upside to using trainable classifiers for records management is the ability to apply retention and sensitivity labels to organise content buried deep within folders or structures. This can help with identifying records and flagging content that may be drafts or duplicates, which can then be deleted. For example, this could include CVs stored outside HR document libraries, or copies and drafts of agreements not stored in a contract management system.

When seeking to use a classifier, it’s always worth first exploring the capabilities of Microsoft pre-trained classifiers. They often fit the bill for recognising and classifying documents.

However, if that’s not the case, and you find that the built-in version isn’t exactly suited to your content, then is it worth investing in a custom trainable classifier?

Pre-trained and custom classifiers

Microsoft has created classifiers that you can start using without the need to train. These classifiers appear with the status of Ready To Use. However, you can also create and train your own custom classifier.

Customisation requires time and effort, so it’s worth examining how much you’ll actually gain from building a custom classifier.

Until recently, there were several steps required to create a custom classifier:

  1. Initial scan of the system
  2. Training the classifier on seed content
  3. Testing the classifier by confirming matches/non-matches with testing documents

Since then, Microsoft has made it easier to create custom trainable classifiers by removing the testing step; training and testing the custom classifier is now one process. You can, however, still test the classifier by uploading additional documents.

Case study: building a custom classifier for contracts

Recently, I created a custom trainable classifier for contracts in response to a client’s request for a demo. Here are the steps I took to achieve this:

1. Gather seed content

The first step in creating my ‘test contracts’ classifier was to gather seed content: at least 50 positive examples of contracts, and 150 negative examples. A colleague pointed me to a developer’s site with publicly available contracts for the positive examples. For the negative examples I used non-sensitive publicly available documents as well as some test documents I created myself. My dataset was initially 100 positive examples and 150 negative examples.

2. Training the classifier on seed content

In Purview, within /Records Management/Classifiers, I chose ‘Create Trainable Classifier’, selecting a SharePoint site and folders for my positive and negative examples. The trainable classifier took a 24-48 hours to be trained, but training was initially unsuccessful, with an accuracy rate of 77%. I added additional documents to the positive sample and removed documents that had been flagged as false positives.

I repeated the process, and this time the accuracy rate was 94.7%.

What I learned about pre-trained vs. custom classifiers

When published to all locations and ready to use, my newly created classifier only tagged 2 contracts (they were correct matches). I also created a retention label for my contracts, and then used an auto-apply policy to publish the label, choosing the trainable classifier as the condition for labelling, and this worked perfectly.

By contrast, Microsoft’s pre-trained classifier ‘Agreements’ tagged all my contract test documents, both seed data and additional test documents, as shown in Content Explorer.

What I learned about trainable classifiers is that using Microsoft pre-trained classifiers may be a better option than creating a custom classifier, where the Microsoft classifier is targeting the same kinds of documents as your custom classifier. Despite having a 94.7% accuracy rate, my custom classifier ‘lost’ consistently in classifying documents to the pre-trained classifier ‘Agreements’. To increase accuracy of my custom classifier, the training dataset has to be larger, which requires a time investment.

On the other hand, a custom classifier I created in the beginning of the year, for a records management ‘file plan’, based on a template, was more successful in identifying matching documents. It’s probable that content that is more unique lends itself to a better custom trainable classifier.

As such, the purpose of the trainable classifier needs to be carefully considered.

Is the classifier meant to identify records and auto-apply a retention label? If there is a risk of false positives (more likely with custom classifiers), applying a label could result in a document being ‘misfiled’ and deleted too early/too late.

Or, is the classifier meant to find drafts/duplicates/downloads of certain key documents in non-record locations, and auto-apply a retention label to delete them after a period of time? As a Records Manager I believe trainable classifiers may be more commonly used in the latter scenario.

In conclusion, Trainable classifiers, exciting as they are, complement other methods of setting retention labels, such as default labels on document libraries, and user filing, rather than replace these.

Receive more blogs like this straight into your inbox

Sign up to receive our latest blogs and stay up to date with our latest news, Microsoft 365 updates, events, webinars and workshops.

Last updated 25 Jun 2025

About the Author: Paula McClure

Paula McClure
A seasoned, expert Information Governance and Privacy professional with 20 years of experience leading Records Management and Data Protection in international organisations and global financial services. Paula is very excited to be stepping into the world of Microsoft 365, helping businesses leverage their information assets in order to enhance decision-making, gain insights, and ensure compliance. Her areas of expertise include financial services regulatory recordkeeping requirements; Records Schedule development; establishing an information governance framework; and implementing file plans on different platforms. Paula has a Master’s in Library and Information Studies (MLIS), is an Accredited member of the Information and Records Management Society and a BCS Practitioner in Data Protection.

Table of contents

Go to Top