> For the complete documentation index, see [llms.txt](https://docs.zingg.ai/latest/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.zingg.ai/latest/stepbystep/createtrainingdata/addowntrainingdata.md).

# Using Pre-existing Training Data

Supplementing Zingg With Existing Training Data

If you already have some training data that you want to start with, you can use that as well with Zingg. Add an attribute **trainingSamples** to the config and define the training pairs.

The training data supplied to Zingg should have a **z\_cluster** column that groups the records together. The **z\_cluster** uniquely identifies the group. We also need to add the **z\_isMatch** column which is **1** if the pairs *match* or **0** if they do *not* match. The **z\_isMatch** value has to be the same for all the records in the **z\_cluster** group. They either match with each other or they don't.

An example is provided in [GitHub training data](https://github.com/zinggAI/zingg/tree/0.7.0/examples/febrl/training.csv).\
Here, the first column specifies the z\_cluster, the second column specifies the z\_isMatch value and the remaining columns are the ones which are used for training the model.

The above training data can be specified using [trainingSamples attribute in the configuration.](https://github.com/zinggAI/zingg/tree/0.7.0/examples/febrl/configWithTrainingSamples.json)

**Note**: It is advisable to still run [findTrainingData](/latest/stepbystep/createtrainingdata/findtrainingdata.md) and [label](/latest/stepbystep/createtrainingdata/label.md) a few rounds to tune Zingg with the supplied training data as well as patterns it needs to learn independently.
