> ## Documentation Index
> Fetch the complete documentation index at: https://docs.cosmosid.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Create a cohort with Query Builder

## What is a cohort?

Building a Cohort outside the Cosmos-Hub typically requires querying metadata spreadsheets, matching that metadata to samples in the Data Table, and a lot of time-consuming data wrangling. On Cosmos-Hub 2.0,  an AI agent will help you assembling the Cohort you need for an Analysis Project, without having to compose a single SQL query.

## What is the Query Builder?

The Query Builder (QB) is an AI-based chatbot that interactively guides you through composing complex metadata queries to build the Data Table used as input for your analyses.

Interacting with the QB is a hybrid experience: you can build your Cohort's metadata query either by accepting filter suggestions in the interface, or by describing what you need in natural language, which the QB interprets and translates into a formal query. As the conversation continues, you refine the query further until it captures the specific Cohort you need.

By the end of the conversation, you will have created an Analysis Project with your queried Cohort, ready for analysis in the Analysis Studio.

## Create a Cohort with the Query Builder

<Steps>
  <Step title="Accessing the Query Builder">
    The Query Builder can be accessed through three different routes

    <Tabs>
      <Tab title="When creating a new Analysis Project">
        <img src="https://mintcdn.com/cmbio/9IPlECfVJPWdo6z4/images/QB_access_1st.jpg?fit=max&auto=format&n=9IPlECfVJPWdo6z4&q=85&s=39fc996cac0898d941a1aa0daf1e8ef0" alt="QB Access 1st" width="2458" height="480" data-path="images/QB_access_1st.jpg" />

        Inside the Study Dashboard, click on "+ New Analysis" to open the Query builder and start the creation of a New Analysis Project.
      </Tab>

      <Tab title="Import Workspace data into a Study">
        1. Click on "Import Data" to open the Data import dialog
        2. Select 'Use Existing Workspace Data' to start building a cohort for the current Study
      </Tab>

      <Tab title="Build a cohort from a data table">
        Typically, the Query builder will first ask *which* type of 'omics profiles (e.g: taxonomic, functional, qPCR, metabolomics) you would like to use as primary source to create a cohort. If you already know which dataset, you can use that dataset as the starting point with the Query Builder by:

        <img src="https://mintcdn.com/cmbio/9IPlECfVJPWdo6z4/images/QB_access_3rd.jpg?fit=max&auto=format&n=9IPlECfVJPWdo6z4&q=85&s=ef47291d45ddd27b00706430eee9b3a1" alt="QB Access 3rd" width="2306" height="1052" data-path="images/QB_access_3rd.jpg" />
      </Tab>
    </Tabs>
  </Step>

  <Step title="Select which dataset(s) will form the basis of the cohort">
    <img src="https://mintcdn.com/cmbio/rXnT_KRi2R3YfV3p/images/qb_1st.jpg?fit=max&auto=format&n=rXnT_KRi2R3YfV3p&q=85&s=d26bb8b49c01b2c6584887bfecf0b8d1" alt="Qb 1st" width="2448" height="508" data-path="images/qb_1st.jpg" />
  </Step>

  <Step title="Add metadata filters using QB suggestions">
    The Query Builder will start by proposing metadata filters based on the sample metadata available for the selected dataset(s):

    * **Categorical Metadata**: as multiple choice options for each metadata attribute, along with the number of samples in that category, and the percentage of the dataset those samples represent.
    * **Numerical Metadata**: as a slider to set the upper and lower limits defining your sample Cohort.

          <img src="https://mintcdn.com/cmbio/rXnT_KRi2R3YfV3p/images/qb_step2.jpeg?fit=max&auto=format&n=rXnT_KRi2R3YfV3p&q=85&s=730f5d1bb67f8dd3ff3cd67101d9d596" alt="Qb Step2" width="2468" height="1216" data-path="images/qb_step2.jpeg" />
  </Step>

  <Step title="Make requests through natural language">
    The QB is a interactive chatbot. You can type any request written in natural language at any point of the conversation. The Query Builder is AI agent built to interpret any written request in natural language and apply the right metadata filters to build the correct query logic.. You can also request  information about the dataset or the current cohort. Try few requests, such as:

    * **Explore the metadata**: Ask which attributes are available and what values they contain. The Query Builder will report which metadata is available for the current sample selection.
          <img src="https://mintcdn.com/cmbio/rXnT_KRi2R3YfV3p/images/qb_step3.jpg?fit=max&auto=format&n=rXnT_KRi2R3YfV3p&q=85&s=27aaf14a983a1016b74c093420c21a78" alt="Qb Step3" width="2462" height="1042" data-path="images/qb_step3.jpg" />

    * **Ask to apply multiple filters at once**

          <img src="https://mintcdn.com/cmbio/rXnT_KRi2R3YfV3p/images/qb_step4.jpg?fit=max&auto=format&n=rXnT_KRi2R3YfV3p&q=85&s=0b21e2e47bb825dba15b5920e1f7f310" alt="Qb Step4" width="2560" height="514" data-path="images/qb_step4.jpg" />
  </Step>

  <Step title="Keep track of metadata filters applied so far">
    As the conversation with the QB goes on, you can keep track of which filters have been applied so far from the Active Filters panel on the right.

    <img src="https://mintcdn.com/cmbio/rXnT_KRi2R3YfV3p/images/squared_filters.jpg?fit=max&auto=format&n=rXnT_KRi2R3YfV3p&q=85&s=3a0537c5150a264466a29766b57d11b4" alt="Squared Filters" width="1002" height="900" data-path="images/squared_filters.jpg" />
  </Step>

  <Step title="Select which Metadata Columns to keep">
    <img src="https://mintcdn.com/cmbio/8lQmi4yfr2V3p2iK/images/columns_to_include.jpg?fit=max&auto=format&n=8lQmi4yfr2V3p2iK&q=85&s=796749ebc092f66abc32f494adab2d56" alt="Columns To Include" width="2284" height="340" data-path="images/columns_to_include.jpg" />

    The Query Builder asks which metadata columns to carry into the Analysis Project. These are the metadata fields that will be available in the Analysis Studio for an Analysis Project

    <Check>
      Not selecting any column will automatically keep all the metadata columns available for the current sample query.
    </Check>
  </Step>

  <Step title="Confirm Creation of the Analysis">
    1. Once you're satisfied with the query built so far, you're ready to finalize the Analysis Project. Click 'Looks good - Create the Analysis'.
    2. Choose a name and a write a small description (optional) for your Analysis Project.
  </Step>

  <Step title="Choose Analyses Modules">
    Choose which Analyses Modules you plan to run on your cohort.

    <img src="https://mintcdn.com/cmbio/9IPlECfVJPWdo6z4/images/Analysis_modules.jpg?fit=max&auto=format&n=9IPlECfVJPWdo6z4&q=85&s=11847f6d7e07fdfc1af081d85339b40e" alt="Analysis Modules" width="2646" height="914" data-path="images/Analysis_modules.jpg" />

    <Check>
      If unsure of which ones you will actually need, click on 'Select All'. This will let you use all the modules available on Cosmos-Hub when analyzing your cohort in the Analysis Studio.
    </Check>
  </Step>

  <Step title="Confirm and Create Analysis">
    The Query Builder will provide a summary of the query built so far. Click on 'Confirm and Create Analysis' to finalize creation of the Analysis Project.

    <img src="https://mintcdn.com/cmbio/rXnT_KRi2R3YfV3p/images/readtocreate.jpg?fit=max&auto=format&n=rXnT_KRi2R3YfV3p&q=85&s=75b248197bc3175a7eb912f8b5fa4b82" alt="Readtocreate" width="2366" height="1306" data-path="images/readtocreate.jpg" />

    <Tip>
      Last minute change of plans? You can adjust your query at any point using the controls at the bottom of the conversation. Continue editing, start over with a new query, or undo the most recently applied filters.
    </Tip>
  </Step>

  <Step title="Track the status of the Analysis Project">
    #### Jobs

    You can check the status of the Analysis Project you just created from the **Jobs** tab in the Study Dashboard. This lets you monitor whether the Analysis Project was created successfully, and if it was not, review the details of the error.

    <img src="https://mintcdn.com/cmbio/IvW3mcdgB0VWEfpX/images/jobs-1.jpg?fit=max&auto=format&n=IvW3mcdgB0VWEfpX&q=85&s=0187f810211e489906a2f96d80b24736" alt="Jobs 1" width="3638" height="600" data-path="images/jobs-1.jpg" />
  </Step>
</Steps>

## Downloading metadata for your current cohort

At any point of the conversation, click **Download Metadata** above the Active Filters panel to inspect and export the metadata for the samples currently selected.

<Frame>
  <img src="https://mintcdn.com/cmbio/IvW3mcdgB0VWEfpX/images/download_metadata.jpg?fit=max&auto=format&n=IvW3mcdgB0VWEfpX&q=85&s=5bdad38859d977336c08452afdecb77e" alt="Download Metadata" width="1456" height="1704" data-path="images/download_metadata.jpg" />
</Frame>

### Column coverage

**Column coverage in this cohort** shows, for each metadata column, the percentage of samples in the cohort that have a value for it. **Export stats as CSV** downloads the coverage figures for every column.

### Choosing which columns to download

The **Columns** section controls what goes into the downloaded file. Filter the list with **Suggested columns**, **All columns** or **Chosen**, then tick columns individually, search by name, or select in bulk with **Suggested**, **All** and **None**. Each row shows that column's coverage, and the counter above the list tracks how many are selected. `sample_id` and `sample_name` are always included.

Click **Download** to export the metadata as a CSV with one row per sample and one column per attribute selected.

<Card title="Analysis Project (AP)" icon="chart-line" href="/analyses/the-analysis-studio">
  Run diversity, differential abundance, ordination and machine learning modules on your cohort.
</Card>
