> For the complete documentation index, see [llms.txt](https://cortex-docs.paloaltonetworks.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://cortex-docs.paloaltonetworks.com/cortex-agentix/configure-cortex-agentix/data-management/dataset-management.md).

# Dataset management

{% hint style="warning" %}

### Prerequisite

Dataset Management requires **View/Edit** RBAC permissions for **Data Management** (under **Configurations** → **Data Management**), which are the same permissions required for Parsing Rules, Data Model Rules, and Event Forwarding.
{% endhint %}

The **Dataset Management** page enables you to manage your datasets and understand your overall data storage duration for different retention periods and datasets based on your hot storage license, and retention add-ons that extend your storage. You can view details about your Cortex AgentiX licenses and retention add-ons by selecting **Settings** → **Cortex AgentiX License**. For more information on license retention and the defaults provided per license, see [Data retention policy](/cortex-agentix/learn-about-cortex-agentix/data-retention-policy.md).

{% hint style="info" %}

### Important

Cortex AgentiX enforces retention on all log-type datasets, excluding metrics and users.
{% endhint %}

<details>

<summary>Hot storage</summary>

Your current hot storage license, including the default license retention and any additional retention add-ons to extend storage, are listed within the **Hot Storage License** section of the **Dataset Management** page. Whenever you extend your license retention, depending on your requirements and license add-ons for hot storage, the add-ons are listed.

</details>

<details>

<summary>Additional hot storage</summary>

You can expand your license retention to include flexible Hot Storage based retention to help accommodate varying storage requirements for different retention periods and datasets. This add-on license is available to purchase based on your storage requirements for a minimum of 1,000 GB. If this license is purchased, an **Additional Storage** subheading in the **Hot Storage License** section is displayed on the **Dataset Management** page with a bar indicating how much of the storage is used.

{% hint style="info" %}

### Note

Only datasets that are already handled as part of the GB license are supported for this license. In addition, the retention configuration is only available in Cortex AgentiX, as opposed to the public APIs.
{% endhint %}

</details>

<details>

<summary>Edit the retention plan</summary>

On any dataset configured to use Additional Hot Storage, you can edit the retention period. This enables you to view the current retention details and configure the retention. This includes setting the amount of flexible hot storage-based retention designated for a dataset and the priority for the dataset's hot storage.

#### How to edit the retention plan

1. Select S**ettings → Configurations → Data Management → Dataset Management**.
2. In the **Datasets** table, right-click any dataset designated with flexible hot storage, and select **Edit Retention Plan**.
3. Set the following parameters:
   * **Additional hot storage**: Set the amount of flexible hot storage-based retention designated for this dataset in months, where a month is calculated as 31 days.
   * **Hot Storage Priority**: Select the priority designated for this dataset's hot storage as either Low, Medium, or High.
4. Click **Save**.

</details>

<details>

<summary>Datasets table</summary>

For each dataset listed in the table, the following information is available:

{% hint style="info" %}

### Note

* Certain fields are exposed and hidden by default. An asterisk (\*) is beside every field that is exposed by default.
* Datasets include dataset permission enforcements in the Cortex Query Language(XQL), Query Center, and XQL Widgets.
  {% endhint %}

| Field                    | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| \*TYPE                   | Displays the type of dataset based on the method used to upload the data. The possible values include: Correlation, Lookup, Raw, Snapshot, System, and User. For more information on each dataset type, see [What are datasets?](/cortex-agentix/configure-cortex-agentix/data-management/dataset-management/what-are-datasets.md).                                                                                                                                                                                      |
| \*LOG UPDATE TYPE        | Event logs are updated either continuously (**Logs**) or the current state is updated periodically (**State**) as detailed in the **Last Updated** column.                                                                                                                                                                                                                                                                                                                                                               |
| \*LAST UPDATED           | <p>Last time the data in the dataset logs were updated.</p><div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Important</strong></p><p>This column is updated once a day. Therefore, if the dataset was created or updated by the target or lookup flows, it's possible that the <strong>Last Updated</strong> value is a day behind when the queries or reports were run as it was before this column was updated.</p></div>                                                 |
| \*ADDITIONAL STORAGE     | Amount of flexible hot storage-based retention designated for this dataset in months, where a month is calculated as 31 days.                                                                                                                                                                                                                                                                                                                                                                                            |
| \*TOTAL DAYS STORED      | Actual number of days that the data is stored in the Cortex AgentiX tenant, which is comprised of the **HOT RANGE**.                                                                                                                                                                                                                                                                                                                                                                                                     |
| \*HOT RANGE              | Details the exact period of the Hot Storage from the start date to the end date.                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| \*TOTAL SIZE STORED      | Actual size of the data that is stored in the Cortex AgentiX tenant. This number is dependent on the events stored in the hot storage. For the **`xdr_data`** dataset, where the first 31 days of storage are included with your license, the first 31 days are not included in the **TOTAL SIZE STORED** number.                                                                                                                                                                                                        |
| \*ADDITIONAL SIZE STORED | Actual size of the additional flexible hot storage data that is stored in the Cortex AgentiX tenant in GB. This number is dependent on the events stored in the hot storage.                                                                                                                                                                                                                                                                                                                                             |
| \*AVERAGE DAILY SIZE     | Average daily amount stored in the Cortex AgentiX tenant. This number is dependent on the events stored in the hot storage.                                                                                                                                                                                                                                                                                                                                                                                              |
| \*HOT STORAGE PRIORITY   | Indicates the priority set for the dataset's hot storage as either **Low**, **Medium**, or **High**.                                                                                                                                                                                                                                                                                                                                                                                                                     |
| \*TOTAL EVENTS           | Number of total events/logs that are stored in the Cortex AgentiX tenant. This number is dependent on the events stored in the hot storage.                                                                                                                                                                                                                                                                                                                                                                              |
| \*AVERAGE EVENT SIZE     | Average size of a single event in the dataset (**TOTAL SIZE STORED** divided by the **TOTAL EVENTS**). This number is dependent on the events stored in the hot storage.                                                                                                                                                                                                                                                                                                                                                 |
| \*TTL                    | For lookup datasets, displays the configured time to live (TTL) before entries expire and are removed automatically. **Forever** means entries never expire. **Custom** uses a set number of days, hours, and minutes. The maximum is 99999 days. For more information, see [Set time to live for lookup datasets](/cortex-agentix/configure-cortex-agentix/data-management/dataset-management/lookup-datasets/set-time-to-live-for-lookup-datasets.md).                                                                 |
| DEFAULT QUERY TARGET     | Details whether the dataset is configured to use as your default query target in XQL Search, so when you write your queries you do not need to define a dataset. By default, only the **`xdr_data`** dataset is configured as the **DEFAULT QUERY TARGET** and this field is set to **Yes**. All other datasets have this field set to **No**. When setting multiple default datasets, your query does not need to mention any of the dataset names, and Cortex AgentiX queries the default datasets using a **`join`**. |
| TOTAL HOT RETENTION      | Total hot storage retention configured for the dataset in months, where a month is calculated as 31 days.                                                                                                                                                                                                                                                                                                                                                                                                                |

</details>

<details>

<summary>Dataset views</summary>

Cortex AgentiX supports creating dataset views in the `Dataset Management` page to enhance data efficiency and security. Dataset views provide a virtual representation of data from one or more datasets, based on the Cortex Query Language (XQL) query defined, and provide multiple benefits, such as joining datasets into logical subsets through defined queries, manipulating data without altering underlying datasets, and segregating data for specific user needs or access privileges through the Role-based access control (RBAC) settings.

Once a dataset view is created, you can edit or delete the dataset view by right-clicking the dataset view in the **Dataset Views** table. A dataset view can only be deleted if there are no other dependencies. For example, if a Correlation Rule is based on a dataset view, you wouldn't be able to delete the dataset view until you removed the dataset view from the XQL query of the Correlation Rule.

Cortex AgentiX logs entries for events related to creating, editing, and deleting datasets or dataset views. These monitored activities are available to view in the datasets and dataset views audit logs in the Management Audit Logs. For more information, see Monitor datasets and dataset views activity.

#### Building XQL dataset view queries

When building an XQL query to define a dataset view, the query is built in the same way as creating a query through the Query Builder. Yet, it's important to be aware of the following points that are specific for dataset view queries:

* The following features are unsupported in dataset view queries:
  * RT Correlation Rules
  * Cortex Data Model (XDM)
  * Query Library
  * Presets
* Only the following XQL stages are supported when building a dataset view query:
  * alter
  * dedup
  * fields
  * filter
  * join
  * replacenull
  * union
* Once the dataset view is created, it is listed as an available `dataset` when building your XQL queries as long as you have the necessary permissions to access the dataset view in the Role-based access control (RBAC) settings.

#### How to create a dataset view

1. Select Settings → Configurations → Data Management → Dataset Management → Dataset Views.
2. Click New Datset View.
3. Enter a Name and Description (optional) for the dataset view.
4. Create your XQL query for the dataset view by typing in the query box.
5. (Optional) Click Run to view the query results.

   The query must contain no errors, including using only supported commands, to run; otherwise, the Run button remain disabled.
6. Click Save.

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><h3>Note</h3><p>You'll only be able to save the dataset view if the query contains no errors; otherwise, the Save button is disabled.</p></div>

Once the dataset view is created, you can now control user access permissions through Role-based access control (RBAC).

#### Dataset views access permissions

{% hint style="info" %}

### Note

License Type: Managing Roles requires an Account Admin or Instance Administrator role.
{% endhint %}

Access permissions for dataset views are configured in the same way that you set dataset access permissions for any dataset through user roles in Cortex AgentiX Access Management. Cortex AgentiX uses role-based access control (RBAC) to manage roles with specific permissions for controlling user access. RBAC helps manage access to Cortex AgentiX components and datasets, so that users, based on their roles, are granted minimal access required to accomplish their tasks. Once the user role is configured to access these dataset views, you can now assign the user role to the designated users or user groups, who you want to access these dataset views.

**How to set access permissions for dataset views**

1. Select Settings → Configurations → Access Management.
2. Configure a user role with the dataset views that you want users to access.
   1. Select Roles.
   2. You can perform one of the following:
      * To create a new role to assign the dataset views, click New Role, and set a Role Name and Description (optional).
      * To edit an existing user role with these dataset views, right-click the relevant user role, and select Edit Role.
      * To create a new role based on an existing role, right-click the relevant user role, select Save As New Role, and set a Role Name and Description (optional).
   3. Under Datasets, you have two options for setting the Cortex Query Language (XQL) dataset access permissions for the user role:
      * Set the user role with access to all XQL datasets by disabling the Enable dataset access management toggle.
      * Set the user role with limited access to certain XQL datasets by selecting the Enable dataset access management toggle and selecting the datasets under the different dataset category headings.
   4. Scroll down to Dataset View and select the particular dataset views that you want assigned to this user role.
   5. Click Save.
3. Assign the user role with the dataset views configured to the designated users or user groups. For more information, see Create a Role in Manage users in the Cortex AgentiX tenant.

#### Dataset Views table

For each dataset view listed in the table, information is available. Here are descriptions on the columns that may require further explanation:

| Field          | Description                                                           |
| -------------- | --------------------------------------------------------------------- |
| SOURCE QUERY   | Displays the query used to create the dataset view.                   |
| IS VALID       | Details whether the query for the dataset view is still valid or not. |
| RELATED TABLES | Details the other datasets that are related to this dataset view.     |

</details>

<details>

<summary>Dataset parsers</summary>

To ensure accurate visibility and threat detection within Cortex AgentiX, the platform utilizes a variety of specialized dataset parsers. The selection and effectiveness of these parsers are directly dependent on the specific data sources ingested, encompassing both the source data (the original log format generated by the device) and the integration configuration. By identifying whether logs arrive as structured JSON, standardized security formats like CEF and LEEF, or unstructured raw text, Cortex AgentiX can properly parse the information for analysis.

**Supported parser types**

* CEF (Common Event Format): A parser for logs formatted in the ArcSight CEF standard. CEF logs contain a pipe-delimited header with vendor, product, version, event class, name, and severity fields, followed by key-value pair extensions. Commonly used by vendors like Check Point, Fortinet, and Zscaler.
* Cisco ASA: A parser for specific Cisco Adaptive Security Appliance (ASA) syslog messages. Supports logs like connection-related messages (Built/Teardown) and AnyConnect VPN events. Only a subset of Cisco ASA message types is supported.
* Corelight: A parser for network traffic logs generated by Corelight sensors (based on Zeek/Bro). Processes structured JSON logs containing network connection metadata such as connection records, DNS queries, and HTTP transactions.
* Filebeat: A parser for logs collected and forwarded using Elastic Filebeat agents. Extracts the log payload from the Filebeat JSON envelope, handling metadata fields such as timestamps, agent information, and host details.
* JSON: A general-purpose parser for logs sent in standard JSON format. This is the most common format for third-party data sources that send structured data, including cloud services, SaaS applications, and API-based integrations.
* LEEF (Log Event Extended Format): A parser for the IBM QRadar LEEF standard. LEEF logs contain a tab-delimited or custom-delimited header with version, vendor, product, product version, and event ID fields, followed by key-value pair attributes. Typically used by IBM security products and QRadar-integrated vendors.

  <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><h3>Note</h3><p>CEF and LEEF logs are often wrapped in a Syslog envelope (RFC 3164/5424). Cortex AgentiX automatically detects and strips the Syslog header to extract the reporting device IP and hostname before parsing the inner CEEF or LEEF payload.</p></div>
* Raw Text: A parser for unstructured, plain-text log data that does not conform to any specific structured format. The raw text is ingested as-is and can be processed using parsing rules. This is also the default parser used when the log format cannot be identified or is not explicitly defined.
* WEC (Windows Event Collection): A parser for Windows Event logs collected via Windows Event Forwarding (WEF/WEC). Parses XML-formatted Windows Event logs and converts them into a structured format. Supports Windows Security, System, and Application event logs.
* Windows DNS Debug: A parser for Microsoft Windows DNS Server debug log files. Extracts DNS query and response details including query type, remote IP, protocol, response code, and question name from the Windows DNS debug log format.
* Winlogbeat: A parser for Windows Event logs collected using Elastic Winlogbeat agents. Similar to the WEC parser but handles the Winlogbeat-specific JSON envelope format, extracting Windows Event data along with associated metadata.

</details>

<details>

<summary>Datasets in XQL</summary>

{% hint style="info" %}

### Important

By default, forensic datasets are not included in XQL query results, unless the dataset query is explicitly defined to use a forensic dataset.
{% endhint %}

Cortex Query Language (XQL) supports using different languages for dataset and field names. In addition, when setting up your XQL query, it is important to keep in mind the following:

* The dataset formats supported are dependent on the data retention offerings available in Cortex AgentiX according to whether you want to query hot storage.
* Dataset refresh times: While most out-of-the-box system datasets are ingested in near real-time, the following datasets have specific refresh schedules.
  * `endpoints`: Refreshed every hour.
  * `pan_dss_raw`: Refreshed daily.
  * Forensics datasets: Data collection behavior depends on your **Agent Settings** profile.
    * Default: Data is collected as a one-time snapshot and does not update.
    * Scheduled: If you specify a collection interval, the value represents the number of hours between updates, such as an interval of 24 equals once per day.
    * Minimum: The shortest allowable interval is 12 hours.
* Query against a dataset by selecting it with the `dataset` command when you create an XQL query. For more information, see [Create XQL query](/cortex-agentix/reference-and-developer-docs/cortex-agentix-xql/build-xql-queries/how-to-build-xql-queries.md).
* After your query runs, you can always save your query results as a dataset. You can use the `target` stage command to save query results as a dataset.
* Schema changes to datasets may not be reflected in the autocomplete suggestions and deﬁnitions as you type in real time the XQL query and can appear with a slight delay.

</details>

<details>

<summary>Managing datasets and dataset views</summary>

You can manage your datasets and dataset views in Cortex AgentiX from the **Settings** → **Configurations** → **Data Management** → **Dataset Management** page.

Below are some of the main tasks available for all dataset types by right-clicking a particular dataset or dataset view listed in either the **Datasets** or **Dataset Views** table. Only tasks that need further explanation are explained below. Datasets and dataset views can only be deleted if there are no other dependencies. For example, if a Correlation Rule is based on a dataset or dataset view or dataset view, you wouldn't be able to delete the dataset or dataset view until you removed the dataset view from the XQL query of the Correlation Rule.

{% hint style="info" %}

### Note

For more information on tasks specific to lookup datasets, see [Lookup datasets](/cortex-agentix/configure-cortex-agentix/data-management/dataset-management/lookup-datasets.md).
{% endhint %}

#### View Schema

Select **View Schema** to view the schema information for every field found in the dataset or dataset view result set in the Schema tab after running the query in XQL. Each system field in the schema is written with an underscore (`_`) before the name of the field in the FIELD NAME column in the table.

{% hint style="info" %}

### Note

Schema changes to datasets may not be reflected in the autocomplete suggestions and deﬁnitions as you type in real time the XQL query and can appear with a slight delay.
{% endhint %}

#### Set as default

Select Set as default to query the dataset without having to specify it in your queries in XQL by typing `dataset = <name of dataset>`. Once configured, the DEFAULT QUERY TARGET column entry for this dataset is set to Yes in the Datasets table. By default, this option is not available when right-clicking the `xdr_data` dataset as this dataset is the only dataset configured as the DEFAULT QUERY TARGET as it contains all of the endpoint and network data that Cortex AgentiX collects. Once you Set as default another dataset, you can always remove it by right-clicking the dataset and selecting Remove from defaults. When setting multiple default datasets, your query does not need to mention any of the dataset names, and Cortex AgentiX queries the default datasets using a `join`. This option is only relevant for datasets.

#### Copy text to clipboard

Select **Copy text to clipboard** to copy the name of the dataset or dataset view to your clipboard.

</details>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://cortex-docs.paloaltonetworks.com/cortex-agentix/configure-cortex-agentix/data-management/dataset-management.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
