# Introduction

Welcome to Extracta LABS, where we are revolutionizing the way businesses interact with data. Our cutting-edge API is designed to streamline the extraction of valuable information from a myriad of document types, making your data processing workflow as efficient and automated as possible.

## Who We Are

Extracta LABS is at the forefront of data extraction technology. We specialize in turning unstructured documents into structured, actionable data. Our team is committed to innovation, quality, and customer satisfaction, ensuring that our solutions are not only powerful but also easy to use and integrate into your existing systems.

## What We Do

Our flagship product, the **API**, is a testament to our dedication to simplifying data extraction. Whether you're dealing with invoices, contracts, resumes, or any other document type, our API provides a robust solution for automatically retrieving the information you need.

## Key Features

* **Versatile Document Processing**: Our API handles various formats, including PDF, Word, TXT, PNG, and JPG, with precision.
* **Customizable Extraction**: Tailor the data extraction process to fit your specific requirements, enhancing the accuracy and relevance of the extracted data.
* **High Accuracy**: Benefit from our advanced algorithms that ensure up to 99% accuracy in data extraction.
* **Ease of Integration**: With our RESTful API, integrating data extraction into your workflow is seamless and straightforward.
* **Data Privacy and Security**: We prioritize the confidentiality and integrity of your data, adhering to the highest data protection standards.

## Our Mission

At Extracta LABS, our mission is to empower businesses by making data extraction not just a necessity but a catalyst for innovation and efficiency. We strive to continuously refine our technology to meet the dynamic needs of our clients, ensuring that every piece of data extracted drives value and insight.

## Our Vision

We envision a world where data is effortlessly accessible, enabling businesses to focus on what truly matters: growth, innovation, and providing value to their customers. By consistently pushing the boundaries of data extraction technology, we aim to be the go-to solution for businesses worldwide.

## Why Extracta LABS?

* **No Pre-training Required**: Our sophisticated algorithm requires no pre-training, making it ready to use right out of the box.
* **User-Defined Templates**: Unlike traditional data extraction tools, we empower users to define their own templates, providing unparalleled flexibility.
* **API-First Service**: Developed with customization and ease of integration in mind, our service is perfect for businesses seeking an adaptable and efficient solution.

## Unlocking Potential

With Extracta's API, the possibilities are endless. From automating workflows to enhancing data accuracy and efficiency, our solution is designed to unlock the full potential of your documents. Join us in redefining the future of data extraction.


# Overview

Unlock the full potential of your document management process with Extracta LABS API — a state-of-the-art solution designed to streamline the extraction of valuable information

## Introduction

Welcome to the Extracta LABS API documentation. Our API is designed to help businesses automate the extraction of data from various document types with high accuracy and efficiency. Whether you're working with invoices, contracts, or any unstructured documents, our API provides the tools you need for seamless integration into your workflow.

## Why Choose Extracta LABS?

* **Versatile Document Processing**: From PDFs to scanned images, our API supports a wide range of document formats.
* **Customizable Extraction**: Tailor the data extraction to fit your specific needs without relying on predefined templates.
* **High Accuracy**: Leverage our fine-tuned algorithms for up to 99% accuracy in data extraction.
* **Data Privacy**: Your data's security is our top priority, with compliance to the highest standards.

## Getting Started

Begin your journey with Extracta LABS by exploring our [Introduction](/) section, which provides essential background information and insights into our API's capabilities. Follow our simple [Authentication](/api-reference/authentication) guide to ensure secure access, and dive into the [Data Extraction - API](/data-extraction-api/api-endpoints-data-extraction) for detailed documentation on each endpoint.

## Explore Our Features

Learn how to make the most of Extracta's API with our [Tutorials](/support/tutorials), which cover everything from basic setup to advanced use cases.

## Stay Connected

Follow us on our social media channels linked in the [Contact Us](/contact/contact-us) section to stay updated on the latest features, updates, and news from Extracta LABS.

<br>


# Authentication

Securing your data and ensuring that access to the Extracta LABS API is protected are our top priorities. This section outlines the process of obtaining an API key, which is necessary for authenticating your API requests.

## Getting Started

To begin using the Extracta LABS API, you'll need to obtain an API key. This key is unique to your account and serves as the credential for accessing the API.

### Step 1: Create an Account

Visit [https://app.extracta.ai](https://app.extracta.ai/) to sign up for an Extracta account. Fill in the required details to register and submit your registration.

### Step 2: Generate an API Key

Once your account is set up, log in and navigate to the `/api` page on the dashboard. Here, you'll find the option to generate a new API key. Click on the "Generate a new API Key" button and follow the prompts. Your new API key will be displayed once generated. Make sure to copy and store it in a secure location.

<figure><img src="/files/W2f7zxB3kpKYjAvPFlCA" alt=""><figcaption></figcaption></figure>

## Using Your API Key

With your API key in hand, you're ready to start making authenticated requests to the Extracta API. To authenticate, include your API key in the header of each request as follows:

```
Authorization: Bearer <Your_API_Key_Here>
```

Replace `<Your_API_Key_Here>` with the API key you generated in the previous step.

## Keeping Your API Key Secure

* **Do not share your API key** publicly or with unauthorized individuals. Treat it as you would your password.
* **Regenerate your API key** if you suspect it has been compromised. You can do this from the same `/api` page where you generated it initially.

## Need Help?

If you encounter any issues while generating or using your API key, please contact support for assistance.

{% content-ref url="/pages/dpJM2AKMBKrLy4SW1M9M" %}
[Contact Us](/contact/contact-us)
{% endcontent-ref %}

<br>


# Supported File Types

At Extracta LABS, we understand the importance of versatility in document extraction processes. Our platform is engineered to accommodate a broad spectrum of document types, ensuring that you can seamlessly integrate our solutions into your workflow, regardless of the document formats you work with.

## Comprehensive Format Support

Our AI-powered extraction technology is designed to handle documents in various formats, including image files, PDFs, and Microsoft Word documents. This capability ensures that Extracta LABS can meet your needs, whether you're processing scanned documents, digital files, or editable documents.

## Currently Supported Formats:

* **Image Files**: Ideal for scanned documents, photographs of documents, and screenshots (.jpeg, .jpg, .png, .tiff, .bmp).
  * ```
    image/jpeg
    ```
  * ```
    image/jpg
    ```
  * ```
    image/png
    ```
  * ```
    image/tiff
    ```
  * ```
    image/bmp
    ```
* **PDF**: Suitable for digital documents that maintain their formatting across different platforms (.pdf).
  * ```
    application/pdf
    ```
* **Microsoft Word Document:** Perfect for editable text documents created in Microsoft Word (.docx, .doc)
  * ```
    application/msword
    ```
  * ```
    application/vnd.openxmlformats-officedocument.wordprocessingml.document
    ```
* **Text Files**: Essential for parsing data from plain text files, facilitating straightforward text extraction without formatting complexities (.txt)
  * ```
    text/plain
    ```

## Processing Capabilities

Extracta's sophisticated document parsing technology not only supports a wide range of file types but also ensures high accuracy in data extraction. Our platform leverages advanced Optical Character Recognition (OCR) techniques to extract text and data efficiently, even from complex document layouts.

## Getting Started

To begin extracting data from your documents, simply upload your files in one of the supported formats via our API. For detailed instructions on how to upload your documents and create extraction requests, please refer to our API documentation.

By supporting multiple document formats, Extracta LABS aims to provide a flexible and comprehensive solution for your document extraction needs. Whether you're working with printed material, digital documents, or editable files, our platform is equipped to deliver precise and reliable extraction results.

{% content-ref url="/pages/Kr8TtNYErc3zJQQ68V7K" %}
[1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction)
{% endcontent-ref %}


# API Endpoints - Data Extraction

Welcome to the Extracta LABS API Reference. This section is your comprehensive guide to the full suite of API endpoints. Our goal is to equip you with all the information needed to integrate our powerful data extraction capabilities seamlessly into your applications and workflows.

## Understanding RESTful APIs

Our API adheres to RESTful principles, making it intuitive and straightforward for developers to use. REST, or Representational State Transfer, is an architectural style that uses HTTP requests to access and manipulate data. The RESTful approach ensures that our API is scalable, reliable, and easy to consume.

## Available Endpoints

The Extracta LABS API offers the following endpoints to facilitate your document parsing needs:

<table data-full-width="false"><thead><tr><th width="271">Endpoint</th><th width="128">Type</th><th>Page</th></tr></thead><tbody><tr><td><code>/createExtraction</code></td><td><mark style="color:green;"><code>POST</code></mark></td><td><a data-mention href="/pages/Kr8TtNYErc3zJQQ68V7K">/pages/Kr8TtNYErc3zJQQ68V7K</a></td></tr><tr><td><code>/viewExtraction</code></td><td><mark style="color:green;"><code>POST</code></mark></td><td><a data-mention href="/pages/MGKeJnCCGwqrMSMnZ1zg">/pages/MGKeJnCCGwqrMSMnZ1zg</a></td></tr><tr><td><code>/updateExtraction</code></td><td><mark style="color:orange;"><code>PATCH</code></mark></td><td><a data-mention href="/pages/5HjxUN1WjPqa8YT7RzBt">/pages/5HjxUN1WjPqa8YT7RzBt</a></td></tr><tr><td><code>/deleteExtraction</code></td><td><mark style="color:red;"><code>DELETE</code></mark></td><td><a data-mention href="/pages/Hfhu9qoavlIQemWr5Wxl">/pages/Hfhu9qoavlIQemWr5Wxl</a></td></tr><tr><td><code>/uploadFiles</code></td><td><mark style="color:green;"><code>POST</code></mark></td><td><a data-mention href="/pages/sWhIeG7Es4VBfiHRfeYw">/pages/sWhIeG7Es4VBfiHRfeYw</a></td></tr><tr><td><code>/getBatchResults</code></td><td><mark style="color:green;"><code>POST</code></mark></td><td><a data-mention href="/pages/n1WqKJmzjygZBzg1bgqZ">/pages/n1WqKJmzjygZBzg1bgqZ</a></td></tr><tr><td><code>/credits</code></td><td><mark style="color:blue;">POST</mark></td><td><a data-mention href="/pages/IkEMD5grdQkC4zfylCDB">/pages/IkEMD5grdQkC4zfylCDB</a></td></tr></tbody></table>

For more details on each endpoint, including request parameters, response objects, and example requests and responses, please navigate to the specific page dedicated to that endpoint.&#x20;

## Postman Collection

For a complete and interactive set of API requests, please refer to our [Postman Integration](/data-extraction-api/postman-integration)collection.

## Navigating the Documentation

For each endpoint, we provide a dedicated page that dives deep into:

* **Request Parameters**: Detailed descriptions of all parameters you can use in your requests, allowing you to customize your API calls to fit your needs precisely.
* **Response Objects**: Insight into the structure and content of the API responses, helping you understand how to interpret and use the data returned by the API.
* **Example Requests and Responses**: Practical examples to guide you through constructing requests and handling responses, making your development process smoother and more efficient.

This structured approach ensures that you have access to all the necessary details to make the most of the Extracta LABS API, regardless of your specific requirements or use case.


# 1. Create extraction

<mark style="color:green;">`POST`</mark> `/createExtraction`

Initiates a new document extraction process. This endpoint allows you to **define** an extraction with specific fields, options, and configurations. Once **created**, you can use the returned `extractionId` to upload files for processing.

## Postman Collection

For a complete and interactive set of API requests, please refer to our [Postman Integration](/data-extraction-api/postman-integration)collection.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="174">Name</th><th width="101">Type</th><th width="104">Required</th><th>Description</th><th>Dedicated page</th></tr></thead><tbody><tr><td><code>name</code></td><td>string</td><td><code>true</code></td><td>A descriptive name for the extraction.</td><td></td></tr><tr><td><code>description</code></td><td>string</td><td><code>false</code></td><td>A description for the extraction.</td><td></td></tr><tr><td><code>language</code></td><td>string</td><td><code>true</code></td><td>Document's language for accurate extraction.</td><td><a data-mention href="/pages/M8nbOGsp9Uqr9YJnlalv">/pages/M8nbOGsp9Uqr9YJnlalv</a></td></tr><tr><td><code>options</code></td><td>object</td><td><code>false</code></td><td>Additional processing options.</td><td><a data-mention href="/pages/BnzkOqj5yUPjNNyPmAEr">/pages/BnzkOqj5yUPjNNyPmAEr</a></td></tr><tr><td><code>fields</code></td><td>object</td><td><code>true</code></td><td>An array of objects, each specifying a field to extract.</td><td><a data-mention href="/pages/N9bjdayNuzaN2PlaDqgm">/pages/N9bjdayNuzaN2PlaDqgm</a></td></tr></tbody></table>

To fully customize your data extraction request, understanding the `fields` parameter is crucial. This parameter allows you to specify exactly what information you want to extract, with options for `string`, `object`, and `array` types to match your data structure needs.

{% content-ref url="/pages/N9bjdayNuzaN2PlaDqgm" %}
[Fields](/data-extraction-api/extraction-details/fields)
{% endcontent-ref %}

Customize your extraction process with additional options such as table analysis and handwritten text recognition.

{% content-ref url="/pages/BnzkOqj5yUPjNNyPmAEr" %}
[Options](/data-extraction-api/extraction-details/options)
{% endcontent-ref %}

## Body Example

```json
{
    "extractionDetails": { 
        "name": "CVs Extraction",
        "description": "...",
        "language": "English",
        "options": {
            "hasTable": false,
            "hasVisuals": false,
            "handwrittenTextRecognition": false,
            "checkboxRecognition": false
        },
        "fields": [
            {
                "description": "",
                "example": "",
                "key": "name"
            },
            {
                "description": "",
                "example": "",
                "key": "surname"
            },
            {
                "description": "",
                "example": "",
                "key": "phone_number"
            },
            {
                "description": "last job title name",
                "example": "Programmer",
                "key": "last_job_position"
            },
            {
                "description": "the number of years in numbers",
                "example": "6",
                "key": "years_of_experience"
            }
        ]
    }
}

```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {
        "name": "CVs Extraction",
        "description": "...",
        "language": "English",
        "options": {
            "hasTable": false,
            "handwrittenTextRecognition": false
        },
        "fields": [
            { "description": "", "example": "", "key": "name" },
            { "description": "", "example": "", "key": "surname" },
            { "description": "", "example": "", "key": "phone_number" },
            { "description": "last job title name", "example": "Programmer", "key": "last_job_position" },
            { "description": "the number of years in numbers", "example": "6", "key": "years_of_experience" }
        ]
    };

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {
        "extractionDetails": {
            "name": "CVs Extraction",
            "description": "...",
            "language": "English",
            "options": {"hasTable": False, "handwrittenTextRecognition": False},
            "fields": [
                {"description": "", "example": "", "key": "name"},
                {"description": "", "example": "", "key": "surname"},
                {"description": "", "example": "", "key": "phone_number"},
                {
                    "description": "last job title name",
                    "example": "Programmer",
                    "key": "last_job_position",
                },
                {
                    "description": "the number of years in numbers",
                    "example": "6",
                    "key": "years_of_experience",
                },
            ],
        }
    }

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param string $token The authorization token for API access.
 * @param array $extractionDetails The details of the extraction to be created.
 * @return mixed The API response with the new extraction ID or an error message.
 */
function createExtraction($token, $extractionDetails) {
    $url = 'https://api.extracta.ai/api/v1/createExtraction';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = json_encode(['extractionDetails' => $extractionDetails]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_POST, 1);

    try {
        // Execute cURL session
        $response = curl_exec($ch);

        // Check for cURL errors
        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        // For simplicity, returning the decoded response for now
        return $response;
    } catch (Exception $e) {
        // Handle exceptions or errors here
        return 'Error: ' . $e->getMessage();
    } finally {
        // Always close the cURL session
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$extractionDetails = [
    "name" => "CVs Extraction",
    "description" => "...",
    "language" => "English",
    "options" => [
        "hasTable" => false,
        "handwrittenTextRecognition" => false
    ],
    "fields" => [
        ["description" => "", "example" => "", "key" => "name"],
        ["description" => "", "example" => "", "key" => "surname"],
        ["description" => "", "example" => "", "key" => "phone_number"],
        ["description" => "last job title name", "example" => "Programmer", "key" => "last_job_position"],
        ["description" => "the number of years in numbers", "example" => "6", "key" => "years_of_experience"]
    ]
];

try {
    $response = createExtraction($token, $extractionDetails);
    echo $response;
} catch (Exception $e) {
    echo "Failed to create new extraction: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "created",
    "createdAt": 1712547789609,
    "extractionId": "extractionId"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Language is required"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Error creating extraction"
}
```

{% endtab %}
{% endtabs %}


# 2. View extraction

<mark style="color:green;">`POST`</mark> `/viewExtraction`

This endpoint retrieves the details of an extraction process previously defined in the system. By submitting the unique `extractionId`, you can obtain information such as the extraction name, language, options set, and the fields that are being extracted. This is useful for verifying the setup of your extraction template or for debugging purposes.

## Postman Collection

For a complete and interactive set of API requests, please refer to our [Postman Integration](/data-extraction-api/postman-integration)collection.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="218">Name</th><th width="126">Type</th><th width="115">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>extractionId</code></td><td>string</td><td><code>true</code></td><td>Unique identifier for the extraction.</td></tr></tbody></table>

## Body Example

```json
{
    "extractionId": "extractionId"
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Retrieves details of an extraction process by its unique extractionId.
 * 
 * @param {string} token - The authorization token to access the API.
 * @param {string} extractionId - The unique identifier for the extraction.
 * @returns {Promise<Object>} The promise that resolves to the extraction details.
 */
async function viewExtraction(token, extractionId) {
    const url = "https://api.extracta.ai/api/v1/viewExtraction";

    try {
        const response = await axios.post(url, {
            extractionId: extractionId
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionId = 'extractionId';

    try {
        const extractionDetails = await viewExtraction(token, extractionId);
        console.log("Extraction Details:", extractionDetails);
    } catch (error) {
        console.error("Failed to retrieve extraction details:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def view_extraction(token, extraction_id):
    """
    Retrieves details of an extraction process by its unique extractionId.

    :param token: The authorization token to access the API.
    :param extraction_id: The unique identifier for the extraction.
    :return: The extraction details as a dictionary.
    """
    
    url = "https://api.extracta.ai/api/v1/viewExtraction"
    
    headers = {
        'Content-Type': 'application/json',
        'Authorization': f'Bearer {token}'
    }
    
    payload = {
        'extractionId': extraction_id
    }

    try:
        response = requests.post(url, json=payload, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles both HTTPError and other request-related errors
        print(f"Failed to retrieve extraction details: {e}")
        return None

# Example usage
if __name__ == "__main__":
    token = 'apiKey'
    extraction_id = 'extractionId'

    extraction_details = view_extraction(token, extraction_id)
    if extraction_details is not None:
        print("Extraction Details:", extraction_details)

```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

/**
 * Retrieves details of an extraction process by its unique extractionId.
 * 
 * @param string $token The authorization token to access the API.
 * @param string $extractionId The unique identifier for the extraction.
 * @return mixed The extraction details or an error message.
 */
function viewExtraction($token, $extractionId) {
    $url = 'https://api.extracta.ai/api/v1/viewExtraction';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = json_encode(['extractionId' => $extractionId]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_POST, 1);

    try {
        // Execute cURL session
        $response = curl_exec($ch);

        // Check for cURL errors
        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        // return the response as a object
        return json_decode($response, true);
    } catch (Exception $e) {
        // Handle exceptions or errors here
        return 'Error: ' . $e->getMessage();
    } finally {
        // Always close the cURL session
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$extractionId = 'extractionId';

try {
    $extractionDetails = viewExtraction($token, $extractionId);
    print_r($extractionDetails);
} catch (Exception $e) {
    echo "Failed to retrieve extraction details: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "extractionId": "extractionId",
    "extractionDetails": {
        "status": "has batches",
        "batches": {
            "33NjeFksJFZVTpLWSFSrlWkxy": {
                "filesNo": 3,
                "origin": "api",
                "startTime": "1699370066649",
                "status": "finished"
            },
            ...
        },
        "name": "API CVs",
        "description": "...",
        "language": "English",
        "options": {
            "handwrittenTextRecognition": true,
            "hasTable": false
        },
        "fields": [
            {
                "description": "",
                "example": "",
                "key": "name"
            },
            ...
        ],
    }
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Extraction does not exist",
    "extractionId": "extractiondId"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Could not get extraction",
    "extractionId": "extractiondId"
}
```

{% endtab %}
{% endtabs %}


# 3. Update extraction

<mark style="color:green;">`PATCH`</mark> `/updateExtraction`

Updates an existing document extraction process by modifying **specified parameters** within the extraction details.&#x20;

Only the parameters **included** in the `extractionDetails` will be **updated**; any parameters **not included** will remain **unchanged**. Use this to efficiently adjust specific aspects of an extraction process without altering its overall configuration.

## Postman Collection

For a complete and interactive set of API requests, please refer to our [Postman Integration](/data-extraction-api/postman-integration)collection.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="183">Name</th><th width="101">Type</th><th width="104">Required</th><th width="199">Description</th><th>Dedicated page</th></tr></thead><tbody><tr><td><code>extractionId</code></td><td>string</td><td><code>true</code></td><td>The extraction Id</td><td></td></tr><tr><td><code>name</code></td><td>string</td><td><code>false</code></td><td>A descriptive name for the extraction.</td><td></td></tr><tr><td><code>description</code></td><td>string</td><td><code>false</code></td><td>A description for the extraction.</td><td></td></tr><tr><td><code>language</code></td><td>string</td><td><code>false</code></td><td>Document's language for accurate extraction.</td><td><a data-mention href="/pages/M8nbOGsp9Uqr9YJnlalv">/pages/M8nbOGsp9Uqr9YJnlalv</a></td></tr><tr><td><code>options</code></td><td>object</td><td><code>false</code></td><td>Additional processing options.</td><td><a data-mention href="/pages/BnzkOqj5yUPjNNyPmAEr">/pages/BnzkOqj5yUPjNNyPmAEr</a></td></tr><tr><td><code>fields</code></td><td>object</td><td><code>false</code></td><td>An array of objects, each specifying a field to extract.</td><td><a data-mention href="/pages/N9bjdayNuzaN2PlaDqgm">/pages/N9bjdayNuzaN2PlaDqgm</a></td></tr></tbody></table>

To fully customize your data extraction request, understanding the `fields` parameter is crucial. This parameter allows you to specify exactly what information you want to extract, with options for `string`, `object`, and `array` types to match your data structure needs.

{% content-ref url="/pages/N9bjdayNuzaN2PlaDqgm" %}
[Fields](/data-extraction-api/extraction-details/fields)
{% endcontent-ref %}

Customize your extraction process with additional options such as table analysis and handwritten text recognition.

{% content-ref url="/pages/BnzkOqj5yUPjNNyPmAEr" %}
[Options](/data-extraction-api/extraction-details/options)
{% endcontent-ref %}

## Body Example

```json
{
    "extractionId": "extractionId",
    "extractionDetails": {
        "name": "CV - English",
        "description": "test",
        "language": "English",
        "options": {
            "hasTable": false,
            "handwrittenTextRecognition": true
        },
        "fields": [
            {
                "key": "name",
                "description": "the name of the person",
                "example": "John"
            },
            {
                "key": "email",
                "description": "the email of the person",
                "example": "john@email.com"
            }
        ]
    }
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Updates an existing document extraction with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {string} extractionId - The ID of the extraction to update.
 * @param {Object} extractionDetails - The new details of the extraction to update.
 * @returns {Promise<Object>} The promise that resolves to the API response with the updated extraction details.
 */
async function updateExtraction(token, extractionId, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/updateExtraction";

    try {
        const response = await axios.patch(url, {
            extractionId,
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionId = 'extractiondId'; // Placeholder for actual extraction ID
    const extractionDetails = {
        "name": "CV - English",
        "description": "...",
        "language": "English",
        "options": {
            "hasTable": false,
            "handwrittenTextRecognition": true
        },
        "fields": [
            {
                "key": "name",
                "description": "the name of the person",
                "example": "John"
            },
            {
                "key": "email",
                "description": "the email of the person",
                "example": "john@email.com"
            }
        ]
    };

    try {
        const response = await updateExtraction(token, extractionId, extractionDetails);
        console.log("Extraction Updated:", response);
    } catch (error) {
        console.error("Failed to update extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def update_extraction(token, extraction_id, extraction_details):
    url = "https://api.extracta.ai/api/v1/updateExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}
    payload = {
        "extractionId": extraction_id,
        "extractionDetails": extraction_details
    }

    try:
        response = requests.patch(url, json=payload, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code indicates an error
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None

# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_id = "extractionId"  # Placeholder for the actual extraction ID to be updated
    extraction_details = {
        "name": "CV - English",
        "description": "...",
        "language": "English",
        "options": {
            "hasTable": False,
            "handwrittenTextRecognition": True
        },
        "fields": [
            {
                "key": "name",
                "description": "the name of the person",
                "example": "John"
            },
            {
                "key": "email",
                "description": "the email of the person",
                "example": "john@email.com"
            }
        ]
    }

    response = update_extraction(token, extraction_id, extraction_details)
    print("Extraction Updated:", response)
```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

/**
 * Updates an existing document extraction with the provided details.
 * 
 * @param string $token The authorization token for API access.
 * @param string $extractionId The ID of the extraction to update.
 * @param array $extractionDetails The new details of the extraction to update.
 * @return mixed The API response with the updated extraction details or an error message.
 */
function updateExtraction($token, $extractionId, $extractionDetails) {
    $url = 'https://api.extracta.ai/api/v1/updateExtraction';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload with both extractionId and extractionDetails
    $payload = json_encode([
        'extractionId' => $extractionId,
        'extractionDetails' => $extractionDetails
    ]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_CUSTOMREQUEST, 'PATCH');
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

    try {
        // Execute cURL session
        $response = curl_exec($ch);

        // Check for cURL errors
        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        // For simplicity, returning the decoded response for now
        return $response;
    } catch (Exception $e) {
        // Handle exceptions or errors here
        return 'Error: ' . $e->getMessage();
    } finally {
        // Always close the cURL session
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$extractionId = 'extractionId'; // Placeholder for the actual extraction ID
$extractionDetails = [
    "name" => "CV - English",
    "description" => "...",
    "language" => "English",
    "options" => [
        "hasTable" => false,
        "handwrittenTextRecognition" => true
    ],
    "fields" => [
        ["key" => "name", "description" => "the name of the person", "example" => "John"],
        ["key" => "email", "description" => "the email of the person", "example" => "john@email.com"]
    ]
];

try {
    $response = updateExtraction($token, $extractionId, $extractionDetails);
    echo $response;
} catch (Exception $e) {
    echo "Failed to update extraction: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "updated",
    "updatedAt": 1712547789609,
    "extractionId": "extractionId"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Extraction does not exist",
    "extractionId": "extractionId"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Could not update extraction",
    "extractionId": "extractionId"
}
```

{% endtab %}
{% endtabs %}


# 4. Delete extraction

<mark style="color:green;">`DELETE`</mark> `/deleteExtraction`

This endpoint enables the deletion of an entire extraction process, a specific batch within an extraction, or an individual file, depending on the parameters provided in the request body. The action is permanent and cannot be undone.

* Providing only the `extractionId` results in the deletion of the entire **extraction** process along with all associated **batches** and **files**.
* Specifying both `extractionId` and `batchId` deletes the specified **batch** and **all files** within it from the extraction.
* Including `extractionId`, `batchId`, and `fileId` leads to the deletion of a **specific** **file** within a batch.

## Postman Collection

For a complete and interactive set of API requests, please refer to our [Postman Integration](/data-extraction-api/postman-integration)collection.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="183">Name</th><th width="135">Type</th><th width="164">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>extractionId</code></td><td>string</td><td><code>true</code></td><td>The extraction id</td></tr><tr><td><code>batchId</code></td><td>string</td><td><code>false</code></td><td>The batch id</td></tr><tr><td><code>fileId</code></td><td>string</td><td><code>false</code></td><td>The file id</td></tr></tbody></table>

## Body Example

```json
{
    "extractionId": "extractionId",
    "batchId": "batchId",
    "fileId": "fileId"
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Deletes a specific file within a batch of an extraction process.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} deletionDetails - The identifiers for the extraction, batch, and file to be deleted.
 * @returns {Promise<Object>} The promise that resolves to the API response confirming deletion.
 */
async function deleteExtraction(token, deletionDetails) {
    const url = "https://api.extracta.ai/api/v1/deleteExtraction";

    try {
        const response = await axios.delete(url, deletionDetails, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const deletionDetails = {
        "extractionId": "yourExtractionId",
        "batchId": "yourBatchId",
        "fileId": "yourFileId"
    };

    try {
        const response = await deleteExtraction(token, deletionDetails);
        console.log(response);
    } catch (error) {
        console.error(error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def delete_extraction(token, deletion_details):
    url = "https://api.extracta.ai/api/v1/deleteExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.delete(url, json=deletion_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code indicates an error
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None

# Example usage
if __name__ == "__main__":
    token = "apiKey"
    deletion_details = {
        "extractionId": "yourExtractionId",
        "batchId": "yourBatchId",
        "fileId": "yourFileId"
    }

    response = delete_extraction(token, deletion_details)
    if response:
        print(response)
    else:
        print("Failed to delete.")
```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

/**
 * Deletes an extraction, a specific batch within an extraction, or a specific file within a batch,
 * depending on the parameters provided.
 * 
 * @param string $token The authorization token for API access.
 * @param array $deletionDetails The identifiers for the extraction, batch, and file to be deleted.
 * @return mixed The API response confirming the deletion or an error message.
 */
function deleteExtraction($token, $deletionDetails) {
    $url = 'https://api.extracta.ai/api/v1/deleteExtraction';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = json_encode($deletionDetails);

    // Set cURL options
    curl_setopt($ch, CURLOPT_CUSTOMREQUEST, 'DELETE');
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

    try {
        // Execute cURL session
        $response = curl_exec($ch);

        // Check for cURL errors
        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        // Optionally, you might want to check the response status code here

        // For simplicity, returning the decoded response for now
        return $response;
    } catch (Exception $e) {
        // Handle exceptions or errors here
        return 'Error: ' . $e->getMessage();
    } finally {
        // Always close the cURL session
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$deletionDetails = [
    "extractionId" => "yourExtractionId",
    "batchId" => "yourBatchId", // Optional, depends on the use case
    "fileId" => "yourFileId" // Optional, depends on the use case
];

try {
    $response = deleteExtraction($token, $deletionDetails);
    echo "Deletion Response: " . $response;
} catch (Exception $e) {
    echo "Failed to delete: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "deleted",
    "deletedAt": 1712547789609
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Extraction id is required"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Could not delete extraction"
}
```

{% endtab %}
{% endtabs %}


# 5. Upload Files

<mark style="color:green;">`POST`</mark> `/uploadFiles`

This endpoint enables users to upload files to a specified extraction. If a `batchId` is included in the request, the files will be added to that specific existing batch on the platform. It is important to ensure that the `batchId` already exists; otherwise, the upload will not be successful. If no `batchId` is provided in the request, a new batch will automatically be created for the files.

Files must be uploaded using the `multipart/form-data` content type, which is suitable for uploading binary files (like documents and images).

## Postman Collection

For a complete and interactive set of API requests, please refer to our [Postman Integration](/data-extraction-api/postman-integration)collection.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value                 |
| ------------- | --------------------- |
| Content-Type  | `multipart/form-data` |
| Authorization | `Bearer <token>`      |

## Body

<table><thead><tr><th width="251">Name</th><th width="119">Type</th><th width="115">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>extractionId</code></td><td>string</td><td><code>true</code></td><td>Unique identifier for the extraction.</td></tr><tr><td><code>batchId</code></td><td>string</td><td><code>false</code></td><td>The ID of the batch to add files to.</td></tr><tr><td><code>files</code></td><td>multipart</td><td><code>true</code></td><td><a data-mention href="/pages/GEp82eoMVXcJO7dA0Z4j">/pages/GEp82eoMVXcJO7dA0Z4j</a></td></tr></tbody></table>

For a seamless extraction process, please ensure your documents are in one of our supported formats. Check our Supported File Types page for a list of all formats we currently accept and additional details to prepare your files accordingly.

{% content-ref url="/pages/GEp82eoMVXcJO7dA0Z4j" %}
[Broken mention](broken://pages/GEp82eoMVXcJO7dA0Z4j)
{% endcontent-ref %}

## Code Example

**Note for PHP Users:** Currently, the `/uploadFiles` endpoint supports uploading only one file per request. Please ensure you submit individual requests for each file you need to upload.

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const fs = require('fs');
const axios = require('axios');
const FormData = require('form-data');

/**
 * Uploads files to the Extracta API using Axios for making HTTP requests.
 * 
 * @param {string} token - The authorization token to access the API.
 * @param {string} extractionId - The ID of the extraction process to which these files belong.
 * @param {Array.<string>} files - Paths to the files to be uploaded.
 * @param {string} [batchId=null] - Optional batch ID if the files belong to a specific batch.
 * @returns {Promise<Object>} The promise that resolves to the API response.
 */
async function uploadFiles(token, extractionId, files, batchId = null) {
    const url = "https://api.extracta.ai/api/v1/uploadFiles";
    let formData = new FormData();

    formData.append('extractionId', extractionId);
    if (batchId) {
        formData.append('batchId', batchId);
    }

    // Append files to formData
    files.forEach(file => {
        formData.append('files', fs.createReadStream(file));
    });

    try {
        const response = await axios.post(url, formData, {
            headers: {
                ...formData.getHeaders(),
                'Authorization': `Bearer ${token}`
            },
            // Axios automatically sets the Content-Type to multipart/form-data with the boundary.
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionId = 'extractionId';
    const files = ['test_1.png', 'test_2.png'];

    try {
        const response = await uploadFiles(token, extractionId, files);
        console.log(response);
    } catch (error) {
        console.error("Failed to upload files:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def upload_files(token, extraction_id, files, batch_id=None):
    url = "https://api.extracta.ai/api/v1/uploadFiles"
    headers = {"Authorization": f"Bearer {token}"}

    # Prepare the files for uploading
    file_streams = [
        (
            "files",
            (
                file,
                open(file, "rb"),
                mimetypes.guess_type(file)[0] or "application/octet-stream",
            ),
        )
        for file in files
    ]
    payload = {"extractionId": extraction_id}
    if batch_id is not None:
        payload["batchId"] = batch_id

    try:
        response = requests.post(url, files=file_streams, data=payload, headers=headers)
        response.raise_for_status()  # This will raise an error for HTTP codes 400 or 500
        return response.json()  # Returns the JSON response if no error
    except requests.HTTPError as e:
        # Print server-side error message
        if response.status_code >= 400:
            error_message = response.json()
            print(f"Server returned an error: {error_message}")
        else:
            print(f"HTTP error occurred: {e}")
    except requests.RequestException as e:
        # Handle other requests exceptions
        print(f"Failed to upload files: {e}")
    except Exception as e:
        # Handle other possible exceptions
        print(f"An unexpected error occurred: {e}")
    return None

# Example usage
if __name__ == "__main__":
    token = 'apiKey'
    extraction_id = 'extractionId'
    files = ['test_1.png', 'test_2.png']

    try:
        response = upload_files(token, extraction_id, files)
        print(response)
    except Exception as e:
        print(f"Failed to upload files: {e}")

```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

/**
 * Uploads files to the Extracta API using cURL for making HTTP requests.
 * 
 * @param string $token The authorization token to access the API.
 * @param string $extractionId The ID of the extraction process to which these files belong.
 * @param string $filePath Path to the file to be uploaded.
 * @param string|null $batchId Optional batch ID if the file belongs to a specific batch.
 * @return mixed The response from the API or an error message.
 */
function uploadFiles($token, $extractionId, $filePath, $batchId = null) {
    $url = 'https://api.extracta.ai/api/v1/uploadFiles';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = [
        'extractionId' => $extractionId,
        'files' => new CURLFile($filePath, 'image/jpeg', basename($filePath))
    ];

    if ($batchId !== null) {
        $payload['batchId'] = $batchId;
    }

    // Set cURL options
    curl_setopt($ch, CURLOPT_POST, 1);
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Authorization: Bearer ' . $token,
        'Content-Type: multipart/form-data'
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

    try {
        // Execute cURL session
        $response = curl_exec($ch);

        // Check for cURL errors
        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        // Here, you could additionally parse the $response if it's in JSON format or if needed,
        // For simplicity, just returning the raw response for now
        return $response;
    } catch (Exception $e) {
        // Handle exceptions or errors here
        return 'Error: ' . $e->getMessage();
    } finally {
        // Always close the cURL session
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$extractionId = 'extractiodId';
$filePath = './test_1.png';
$batchId = null;

try {
    $response = uploadFiles($token, $extractionId, $filePath, $batchId);
    echo $response;
} catch (Exception $e) {
    echo "Failed to upload file: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "uploaded",
    "extractionId": "extractionId",
    "batchId": "batchId",
    "files": [
        {
            "fileId": "fileId",
            "fileName": "fileName",
            "numberOfPages": 1,
            "url": "url"
        }
    ]
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Extraction does not exist",
    "extractionId": "extractionId"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Error uploading files"
}
```

{% endtab %}
{% endtabs %}


# 6. Get results

<mark style="color:green;">`POST`</mark> `/getBatchResults`

This endpoint retrieves the results for a specific batch of documents. By providing the `extractionId` and `batchId`, you can obtain the processed data or the current status of the batch, indicating whether the processing is complete or still in progress. When also providing the `fileId`, the endpoint filters the results to only that file.

One alternative will be to use the [Webhook](/data-extraction-api/webhook) to receive the data when its ready on your own server.

## Postman Collection

For a complete and interactive set of API requests, please refer to our [Postman Integration](/data-extraction-api/postman-integration)collection.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="242">Name</th><th width="99">Type</th><th width="126">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>extractionId</code></td><td>string</td><td><code>true</code></td><td>The ID of the extraction</td></tr><tr><td><code>batchId</code></td><td>string</td><td><code>true</code></td><td>The ID of the batch</td></tr><tr><td><code>fileId</code></td><td>string</td><td><code>false</code></td><td>The ID of the file</td></tr></tbody></table>

## Body Example

```json
{
    "extractionId": "extractionId",
    "batchId": "batchId",
    "fileId": "fileId" // optional
}
```

## ⚠️ Important

To avoid rate-limiting, please ensure a delay of 2 seconds between consecutive requests to this endpoint.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Retrieves the results for a specific batch of documents.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {string} extractionId - The unique identifier for the extraction.
 * @param {string} batchId - The unique identifier for the batch.
 * @param {string} [fileId] - The unique identifier for the file (optional).
 * @returns {Promise<Object>} The promise that resolves to the batch results.
 */
async function getBatchResults(token, extractionId, batchId, fileId) {
    const url = "https://api.extracta.ai/api/v1/getBatchResults";

    try {
        // Constructing the request payload
        const payload = {
            extractionId,
            batchId
        };

        // Adding fileId to the payload if provided
        if (fileId) {
            payload.fileId = fileId;
        }

        const response = await axios.post(url, payload, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionId = 'extractionId';
    const batchId = 'batchId';
    const fileId = 'optionalFileId'; // Set this to null or undefined if you don't want to include it

    try {
        const batchResults = await getBatchResults(token, extractionId, batchId, fileId);
        console.log("Batch Results:", batchResults);
    } catch (error) {
        console.error("Failed to retrieve batch results:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def get_batch_results(token, extraction_id, batch_id, file_id=None):
    """
    Retrieves the results for a specific batch of documents.

    :param token: The authorization token for API access.
    :param extraction_id: The unique identifier for the extraction.
    :param batch_id: The unique identifier for the batch.
    :param file_id: The unique identifier for the file (optional).
    :return: The batch results as a Python dictionary.
    """
    url = "https://api.extracta.ai/api/v1/getBatchResults"
    
    headers = {
        'Content-Type': 'application/json',
        'Authorization': f'Bearer {token}'
    }
    
    payload = {
        'extractionId': extraction_id,
        'batchId': batch_id
    }
    
    # Adding file_id to the payload if provided
    if file_id:
        payload['fileId'] = file_id

    response = requests.post(url, json=payload, headers=headers)
    
    # Check if the request was successful
    if response.status_code == 200:
        return response.json()  # Return the parsed JSON response
    else:
        # Handle errors or unsuccessful responses
        response.raise_for_status()

# Example usage
if __name__ == "__main__":
    token = 'apiKey'
    extractionId = 'extractionId'
    batchId = 'batchId'
    fileId = 'optionalFileId'  # Set this to None if you don't want to include it

    try:
        batch_results = get_batch_results(token, extractionId, batchId, fileId)
        print("Batch Results:", batch_results)
    except Exception as e:
        print("Failed to retrieve batch results:", e)

```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

/**
 * Retrieves the results for a specific batch of documents.
 * 
 * @param string $token The authorization token for API access.
 * @param string $extractionId The unique identifier for the extraction.
 * @param string $batchId The unique identifier for the batch.
 * @param string|null $fileId The unique identifier for the file (optional).
 * @return mixed The batch results or an error message.
 */
function getBatchResults($token, $extractionId, $batchId, $fileId = null) {
    $url = 'https://api.extracta.ai/api/v1/getBatchResults';

    // Initialize cURL session
    $ch = curl_init($url);
    
    // Prepare the payload
    $payload = [
        'extractionId' => $extractionId,
        'batchId' => $batchId
    ];

    // Adding fileId to the payload if provided
    if ($fileId !== null) {
        $payload['fileId'] = $fileId;
    }

    $payload = json_encode($payload);

    // Set cURL options
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_POST, 1);

    try {
        // Execute cURL session
        $response = curl_exec($ch);

        // Check for cURL errors
        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        return $response;
    } catch (Exception $e) {
        // Handle exceptions or errors here
        return 'Error: ' . $e->getMessage();
    } finally {
        // Always close the cURL session
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$extractionId = 'extractionId';
$batchId = 'batchId';
$fileId = 'optionalFileId'; // Set this to null if you don't want to include it

try {
    $batchResults = getBatchResults($token, $extractionId, $batchId, $fileId);
    echo $batchResults;
} catch (Exception $e) {
    echo "Failed to retrieve batch results: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200 for all files" %}

```json
{
    "extractionId": "extractionId",
    "batchId": "batchId",
    "files": [
        {
            "fileName": "File 2.png",
            "status": "processed",
            "result": {
                "last_job_position": "Full-Stack Developer",
                "name": "John",
                "phone_number": "000 000 000",
                "surname": "Smith",
                "years_of_experience": "6"
            },
            "url": "fileUrl"
        },
        ...
    ]
}
```

{% endtab %}

{% tab title="200 for a single file" %}

```json
{
    "extractionId": "extractionId",
    "batchId": "batchId",
    "fileId": "fileId",
    "files": [
        {
            "fileName": "File 1.png",
            "status": "processed",
            "result": {
                "last_job_position": "Full-Stack Developer",
                "name": "John",
                "phone_number": "000 000 000",
                "surname": "Smith",
                "years_of_experience": "6"
            },
            "url": "fileUrl"
        }
    ]
}
```

{% endtab %}

{% tab title="200 status waiting" %}

```json
{
    "extractionId": "extractionId",
    "batchId": "batchId",
    "fileId": "fileId", // optional
    "status": "waiting"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "error": "Invalid request"
}
```

{% endtab %}
{% endtabs %}


# 7. Get credits

<mark style="color:blue;">GET</mark> `/credits`

This GET endpoint returns the current number of credits available on your account. Credits are consumed per page processed, where **1 credit = 1 page**.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "ok",
    "credits": 50 // pages
}
```

{% endtab %}
{% endtabs %}


# Extraction Details


# Supported Languages

Extracta LABS is dedicated to breaking language barriers in document extraction, offering robust support for a diverse range of languages. Our advanced AI algorithms are designed to accurately parse and extract information from documents in multiple languages, ensuring that businesses worldwide can leverage our technology.

## Language Parameter

When creating an extraction request, specify the language of your document using the `language` parameter in the extraction details. This helps tailor the extraction process to the linguistic nuances of the content, enhancing accuracy and efficiency.

```json
{
  "extractionDetails": {
    "language": "Multi-Lingual"
  }
}
```

## Currently Supported Languages

Below is a list of languages currently supported by Extracta LABS. Use the corresponding string when specifying the language in your extraction requests.

Arabic, Bangla, Bulgarian, Croatian, Czech, English, Filipino, French, German, Hindi, Hungarian, Italian, Nepali, Polish, Portuguese, Romanian, Russian, Serbian, Spanish, Turkish, Ukrainian, Urdu, and Vietnamese.

#### We are also excited to announce the addition of **'Multi-Lingual'** support. This new feature allows for the processing of documents in any language, ensuring that Extracta LABS can handle your documents regardless of the language used. This enhancement makes our platform even more versatile and accommodating to diverse linguistic needs.

<table><thead><tr><th>Language</th><th width="225"></th><th></th></tr></thead><tbody><tr><td>Arabic</td><td>German</td><td>Russian</td></tr><tr><td>Bangla</td><td>Hindi</td><td>Serbian</td></tr><tr><td>Bulgarian</td><td>Hungarian</td><td>Spanish</td></tr><tr><td>Croatian</td><td>Italian</td><td>Turkish</td></tr><tr><td>Czech</td><td>Nepali</td><td>Ukrainian</td></tr><tr><td>English</td><td>Polish</td><td>Urdu</td></tr><tr><td>Filipino</td><td>Portuguese</td><td>Vietnamese</td></tr><tr><td>French</td><td>Romanian</td><td><strong>Multi-Lingual</strong></td></tr></tbody></table>

{% content-ref url="/pages/Kr8TtNYErc3zJQQ68V7K" %}
[1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction)
{% endcontent-ref %}


# Options

The `options` object in Extracta's API requests provides additional customization for your data extraction process. By setting different properties within this object, you can tailor the extraction to suit the specific needs of your documents. Here's how each option affects your data extraction:

## Options Overview

### `hasTable`

* **Type**: Boolean
* **Not required**
* **Default**: `false`
* **Description**: Indicates whether the document to be processed contains tables. When set to `true`, the extraction process includes an additional step specifically designed to analyze and extract information from tables within the document. This option ensures that table data is accurately recognized and extracted, providing structured information that's easy to use and analyze.

### `hasVisuals`

* **Type**: Boolean
* **Not required**
* **Default**: `false`
* **Description**: Indicates whether the document contains visual elements such as charts, graphs and diagrams that need to be processed. When set to true, the extraction process includes an additional step designed to detect, analyze and extract structured data from these visual components.

### `handwrittenTextRecognition`

* **Type**: Boolean
* **Not required**
* **Default**: `false`
* **Description**: Determines if the document includes handwritten text that needs to be recognized and extracted. Setting this option to `true` initiates a specialized step in the extraction process focused on analyzing handwritten text. This feature leverages advanced OCR and machine learning techniques to convert handwritten notes into digital text, enhancing the comprehensiveness of the data extraction.

### `checkboxRecognition`

* **Type**: Boolean
* **Not required**
* **Default**: `false`
* **Description**: Determines if the document contains checkboxes that need to be recognized and their states (checked or unchecked) extracted. When set to true, the extraction process includes a specialized step focused on identifying checkboxes within the document and accurately determining their status.

### `longDocument`

* **Type**: Boolean
* **Not required**
* **Default**: `false`
* **Description**: Designed for handling very large or complex documents that may exceed typical processing limits. When set to true, the extraction process applies optimized methods to ensure stability and accuracy across long documents, reducing the chance of timeouts or incomplete results. This option is particularly useful for reports, research papers, legal contracts, or any file with extensive text content.

### `splitPdfPages`

* **Type**: Boolean
* **Not required**
* **Default**: `false`
* **Description**: Controls whether a PDF is split into individual pages before processing. When set to true, each page is treated as a separate extraction unit, improving performance and accuracy for multi-page PDFs. This option is ideal when you need fine-grained control over page-level data or want to parallelize processing for faster results. Think of a document with 100 pages full of one-page invoices, in this case you want to extract data for every page, so set the parameter to true.

### `specificPageProcessing`

* **Type**: Boolean
* **Not required**
* **Default**: `false`
* **Description**: Specific Page Processing is a feature designed to allow users to extract and process only a specified range of pages from a PDF document rather than processing the entire document. This feature is particularly useful when working with large PDF files where only certain sections or pages are relevant for the task at hand.

### `specificPageProcessingOptions`

* **Type**: Map
* **Required only if** `specificPageProcessing` **is** `true`
* **Description**: When **Specific Page Processing** is enabled, the system allows the user to define a range of pages (using `from` and `to` parameters) that they want to focus on. The specified range of pages is then extracted from the original PDF, creating a new document that contains only those pages. This newly created PDF is treated as a separate file and is processed according to the usual workflow—whether for storage, analysis, or further manipulation.
* **Example:** Imagine a scenario where you have a 100-page PDF document, but the relevant information that you want to extract is from page 1 to 3. This feature allows you to specify this range, reducing processing time and cost.

<pre class="language-json"><code class="lang-json"><strong>{
</strong>  "options": {
    "specificPageProcessing": true,
    "specificPageProcessingOptions": {
      "from": 1,
      "to": 3
    }
  }
}
</code></pre>

## Using the `options` Object

To utilize these options, include the `options` object in your API request payload, specifying your preferences for `hasTable` and `handwrittenTextRecognition` as shown below:

<pre class="language-json"><code class="lang-json"><strong>{
</strong>  "options": {
    "hasTable": false,
    "handwrittenTextRecognition": false,
    "checkboxRecognition": false
  }
}
</code></pre>

Adjust the values according to the needs of your document. For instance, if your document includes **tables** and **handwritten** **notes**, your `options` object would look like this:

```json
{
  "options": {
    "hasTable": true,
    "handwrittenTextRecognition": true,
    "checkboxRecognition": false
  }
}
```

## Conclusion

The `options` object allows for significant customization of the extraction process, enabling you to adapt the extraction to fit the unique characteristics of your documents. Whether dealing with complex tables, handwritten notes, or both, adjusting these options ensures that your extraction process is optimized for the highest accuracy and relevance of the extracted data.

Remember to review your documents' needs and set the `options` accordingly to take full advantage of the customized extraction capabilities offered by Extracta LABS.

{% content-ref url="/pages/Kr8TtNYErc3zJQQ68V7K" %}
[1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction)
{% endcontent-ref %}


# Fields

The `fields` parameter plays a crucial role in defining the data you wish to extract using the Extracta LABS API. It allows you to specify the exact structure of the information you're interested in, ensuring tailored and efficient data extraction. There are three primary types of fields you can define: `string`, `object`, and `array`. Understanding the differences and applications of each is key to leveraging the Extracta LABS API effectively.

## String Type

The `string` type is the simplest form and is used to extract text-based information. Each `string` field requires a `key`, `description`, `example`, and a `type`.

**Example:**

```json
{
  "key": "name",
  "description": "Name of the person",
  "example": "Alex Smith",
  "type": "string"
}
```

This defines a single piece of data you wish to extract, such as a name or email address, where the value is expected to be a text string.

## Object Type

The `object` type is used for structured data that includes multiple related properties. An `object` field will have a `key`, `description`, `type`, and a list of `properties`, each of which is itself a field definition that can be a `string`, another `object`, or an `array`.

**Example:**

```json
{
  "key": "personal_info",
  "description": "Personal information of the person",
  "type": "object",
  "properties": [
    {
      "key": "name",
      "description": "Name of the person",
      "example": "Alex Smith",
      "type": "string"
    },
    {
      "key": "email",
      "description": "Email of the person",
      "example": "alex.smith@gmail.com",
      "type": "string"
    }
  ]
}
```

This structure is ideal for extracting grouped data, such as personal information, that contains multiple attributes.

## Array Type

The `array` type is used when the data to be extracted is a list of items, which can be either simple `string` values or structured `object` types. An `array` field will include a `key`, `description`, `type`, and `items` specifying the type of elements in the array.

### Array of Strings Example:

```json
{
  "key": "languages",
  "description": "Languages spoken by the person",
  "type": "array",
  "items": {
    "type": "string",
    "example": "English"
  }
}
```

This format is used for lists where each item is a text string, like languages or skills.

### Array of Objects Example:

```json
{
  "key": "items",
  "description": "The items in the invoice",
  "type": "array",
  "items": {
    "type": "object",
    "properties": [
      {
        "key": "name",
        "description": "The name of the item",
        "example": "Item 1",
        "type": "string"
      },
      {
        "key": "quantity",
        "description": "The quantity of the item",
        "example": "1",
        "type": "string"
      },
      {
        "key": "unit_price",
        "description": "The unit price of the item. Return only the number as a string.",
        "example": "100.00",
        "type": "string"
      }
    ]
  }
}
```

This structure supports extracting a list of complex items, each with its own set of attributes, such as invoice items.

## Conclusion

Understanding the distinction between `string`, `object`, and `array` types is fundamental when defining the `fields` parameter for your data extraction needs with Extracta LABS. By carefully structuring your `fields`, you can customize the API's output to match the specific requirements of your application, ensuring that you capture precisely the data you need.

{% content-ref url="/pages/Kr8TtNYErc3zJQQ68V7K" %}
[1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction)
{% endcontent-ref %}


# Polling vs Webhook

## Polling

Polling is a method where your application will repeatedly make a request to the `/getBatchResult` endpoint to check if the batch processing has been completed. This is done by sending requests at regular intervals, using the `extractionId` and `batchId` received from the `/uploadFiles` endpoint.

### **Instructions for Polling:**

1. After calling `/uploadFiles`, store the `extractionId` and `batchId` from the response.
2. Make a `POST` request to `/getBatchResult` with the `extractionId` and `batchId`.
3. If the batch is not yet processed, wait a few seconds and then send the request again.
4. Repeat step 3 until you receive a response indicating that the batch status is `finished`.

***

## Webhook

Webhooks provide a way to receive a callback notification to a specified URL endpoint when an event occurs, in this case, when a batch processing is complete. Instead of your application checking in at regular intervals, the API will send a POST request to the endpoint you configured with the results once they are ready.

### **Instructions for Using Webhook:**

1. Configure your webhook URL and secret in the dashboard.
2. Set up an HTTP server with an endpoint to listen for POST requests from the Extracta's API.
3. Validate the incoming requests using the provided signature in the headers to ensure they are from Extracta LABS.
4. Once validated, process the data sent by the webhook as needed.

Using webhooks is generally more efficient than polling, as it eliminates the need for repeated requests and provides real-time updates as soon as the batch processing is complete. However, setting up a webhook requires you to have a publicly accessible URL and handle the security for verifying incoming requests.


# Webhook (LEGACY)

## **Legacy Webhook Notice**

You are currently using an older version of the Extracta LABS webhook listener.

This version is still **functional**, but it does **not support newer features** such as event types and structured payload handling ([Webhook Event Types](/data-extraction-api/webhook-event-types)). To access the latest improvements, we recommend upgrading your webhook integration.

👉 Visit the **API page** in your **Extracta** **LABS dashboard** to update your webhook listener to the latest version. Upgrading ensures you get the most accurate, flexible, and future-proof webhook experience.

This is the new webhook: [Webhook](/data-extraction-api/webhook)

***

## Description

Webhooks allow you to receive real-time notifications of events happening within your Extracta  LABS extractions. This section will guide you through setting up a Node.js server with Express to listen for and handle webhook events securely.

***

## Prerequisites

* Node.js installed on your server
* An Express.js application
* A secret key obtained from the Extracta LABS dashboard

***

## Step 1: Set Up Your Server

First, ensure you have Express and the necessary packages installed in your project. If not, you can install them using npm:

```bash
npm install express body-parser crypto --save
```

***

## Step 2: Implement Webhook Endpoint

Create a basic HTTP server with Express to listen for webhook POST requests. Use the following code snippet as a starting point:

{% code title="server.js" lineNumbers="true" fullWidth="false" %}

```javascript
const express = require('express');
const crypto = require('crypto');
const bodyParser = require('body-parser');

const app = express();
const port = 4000;

app.use(bodyParser.json());

// Your webhook secret key from the dashboard
const secret = 'secretKey';

// Middleware to validate the webhook signature
function validateSignature(req, res, next) {
  const sigHeader = req.headers["x-webhook-signature"];

  const signature = crypto.createHmac('sha256', secret.replace('E_AI_K_', '')).update(req.body.result).digest('base64');

  if (signature !== sigHeader) {
    return res.status(401).send({ message: "Webhook is not properly signed" })
  }

  next()
}

// Webhook endpoint
app.post('/webhook', validateSignature, (req, res) => {
  console.log('Webhook received:', req.body);
  
  // Process the webhook payload as needed
  // ...

  res.send({ message: "Webhook received" })
})

app.listen(port, () => console.log(`Server listening on port ${port}!`))
```

{% endcode %}

***

## Step 3: Test Your Webhook Listener

After setting up your webhook listener, test it by sending a simulated files. Ensure your server validates the signature and processes the event correctly.

***

By following these steps, you can securely set up your application to receive and process webhook events from Extracta LABS, enabling real-time updates and actions based on the events transmitted to your endpoint.


# Webhook

Webhooks allow you to receive real-time notifications of events happening within your Extracta LABS extractions. This section will guide you through setting up a Node.js server with Express to securely listen for and handle webhook events.

***

## Webhook Payload Structure

Each webhook sent by Extracta LABS will include two primary fields in the request body:

* `event` – A string identifying the type of event (e.g., `extraction.processed`, `extraction.failed`).
* `result` – A list containing files data

Example payload:

```json
{
  "event": "extraction.processed",
  "result": [
    {
      "extractionId": "extractionId",
      "batchId": "batchId",
      "fileId": "fileId",
      "fileName": "fileName",
      "status": "processed",
      "result": {},
      "url": "fileUrl"
    }
  ]
}
```

***

## See all event types

{% content-ref url="/pages/mknkVWnaDyxToEeyPBY6" %}
[Webhook Event Types](/data-extraction-api/webhook-event-types)
{% endcontent-ref %}

***

## Prerequisites

* Node.js installed on your server
* An Express.js application
* A secret key obtained from the Extracta LABS dashboard

***

## Step 1: Set Up Your Server

First, ensure you have Express and the necessary packages installed in your project. If not, you can install them using npm:

```bash
npm install express body-parser crypto --save
```

***

## Step 2: Implement Webhook Endpoint

Create a basic HTTP server with Express to listen for webhook POST requests. Use the following code snippet as a starting point:

{% code title="server.js" fullWidth="false" %}

```javascript
const express = require('express');
const crypto = require('crypto');
const bodyParser = require('body-parser');

const app = express();
const port = 4000;

app.use(bodyParser.json());

// Your webhook secret key from the dashboard
const secret = 'secretKey';

// Middleware to validate the webhook signature
function validateSignature(req, res, next) {
    const signatureRequest = req.headers["x-webhook-signature"];

    const resultString = JSON.stringify(req.body.result);
    const signature = crypto.createHmac('sha256', secret.replace('E_AI_K_', '')).update(resultString).digest('base64');

    if (signature !== signatureRequest) {
        return res.status(401).send({
            message: "Webhook is not properly signed"
        });
    }

    return next();
}

// Webhook endpoint
app.post('/webhook', validateSignature, async (req, res) => {
    try {
        let { event, result } = req.body;

        switch (event) {
            case "extraction.processed":
                console.log("extraction.processed", result);
                break;
            case "extraction.edited":
                console.log("extraction.edited", result);
                break;
            case "extraction.confirmed":
                console.log("extraction.confirmed", result);
                break;
            case "extraction.failed":
                console.log("extraction.failed", result);
                break;
            default:
                console.log("unknown event", event);
                break;
        }

        return res.send({
            event: event,
            message: "Webhook received",
            timestamp: new Date().toISOString()
        })
    } catch (error) {
        console.error("Error processing webhook", error);

        return res.status(500).send({
            message: "Error processing webhook",
            error: error.message,
            timestamp: new Date().toISOString()
        });
    }
})

app.listen(port, () => console.log(`Server listening on port ${port}!`))
```

{% endcode %}

***

## Step 3: Test Your Webhook Listener

Once your webhook listener is set up, test it by triggering events from Extracta LABS. Confirm that:

* The signature is validated correctly.
* The `event` is identified.
* The `result` is handled based on the event type.

***

By following these steps, you can securely set up your application to receive and process webhook events from Extracta LABS, enabling real-time updates and actions based on the events transmitted to your endpoint.


# Webhook Event Types

Extracta LABS webhooks send an `event` field in the payload to indicate what type of action occurred in the system. The event name is a **namespaced string**, with the format:

```
<feature>.<action>
```

This makes the system **modular and scalable**, allowing multiple features (e.g., data extraction, document classification, etc.) to have their own distinct set of events without overlap.

***

## 🔍 Current Feature: `extraction`

All current webhook events are prefixed with `extraction.` because they relate to **data extraction**, which is the first core feature of Extracta LABS.

***

## ✅ Available `extraction` Events

### **`extraction.processed`**

**Description:**\
Triggered when one or more files are successfully processed and data has been extracted. If you process multiple files in a single upload, this webhook will be triggered when all the files are processed.

**Why it's useful:**\
Initiate follow-up actions like storing results, analyzing data, or notifying stakeholders.

***

### **`extraction.edited`**

**Description:**\
Fired when a user manually edits the extracted data in the Extracta LABS interface.

**Why it's useful:**\
Track user input, sync updates to downstream systems, or trigger quality checks.

***

### **`extraction.confirmed`**

**Description:**\
Occurs when a user manually confirms the extracted data, indicating final approval.

**Why it's useful:**\
Use this as a signal to finalize workflows, store results permanently, or notify teams.

***

### **`extraction.failed`**

**Description:**\
Sent when a file fails to process (e.g. due to corruption, unsupported format, or internal errors).

**Why it's useful:**\
Alert your team, retry processing, or prompt the user to upload a new version.

***

## 🧠 Future-Proofing: Other Feature Namespaces

As Extracta LABS grows, additional webhook namespaces will be introduced. For example:

**`classification.*` (future)**

***

By using namespaced events like `extraction.processed`, your application can easily distinguish between different types of events and scale accordingly as new features are added.


# Postman Integration

<div align="left"><figure><img src="/files/y2yBc3LbuDb3UnlwB0pz" alt=""><figcaption></figcaption></figure></div>

To help you get started quickly with the Extracta LABS API, we have created a comprehensive Postman collection. This collection includes all the API endpoints, complete with example requests and responses, making it easier for you to test and integrate with our API.

### How to Use the Postman Collection

* **Import the Collection into your Postman account:** <https://www.postman.com/extracta/workspace/extracta-ai-workspace/collection/35309516-6de7ea3c-857e-486e-b721-7488189881bf?action=share&creator=35309516>
* **Set Up Your Environment:**
  * Configure your variables in **Collection -> Variables** to include your **API key**.
  * How to make an API Key: [Authentication](/api-reference/authentication)

<figure><img src="/files/AqwQvYSrAWlm9QcRO2YW" alt=""><figcaption></figcaption></figure>

* **Explore the API Endpoints:** You can easily navigate through the endpoints and see example requests and responses.

### Import directly the collection by JSON File.

* **Download JSON File with API Endpoints**: <https://drive.google.com/file/d/1PLK1a2-LPr30XBfEdJYaCNmiFx34J-PU/view?usp=sharing>
* **Import the file in Postman as in the screenshot below:**

<figure><img src="/files/oaTqdKlJ85dMAK9RY1d4" alt=""><figcaption></figcaption></figure>

### Benefits of Using the Postman Collection

* **Quick Start:** Instantly access all available API endpoints with pre-configured requests.
* **Ease of Use:** Easily test different endpoints and see example responses without writing any code.
* **Consistency:** Ensure that you are using the correct request formats and parameters as per our latest API specifications.

### Additional Resources

* **API Endpoints Page:** [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction)
* **Support:** If you encounter any issues or have questions, please visit [Contact Us](/contact/contact-us) page.

We hope this Postman collection makes your development process smoother and more efficient. Happy coding!


# API Endpoints - Document Classification

Welcome to the Extracta LABS **Document Classification API**. This section guides you through all available endpoints designed to help you classify uploaded documents into user-defined types and optionally extract structured data from them.

Whether you're automating invoice routing, separating receipts from contracts, or identifying CVs, this API gives you precise control over classification and downstream processing.

## What is Document Classification?

Document Classification is the process of automatically assigning a document to a predefined category — such as an Invoice, Receipt, or Purchase Order — based on its content.

You define the categories (called [Document Types](/document-classification-api/classification-details/document-types)), provide keywords, and optionally link them to extraction templates. Once classified, the document can automatically be parsed for structured data.

## 🔌 Available Endpoints

The Extracta LABS Classification API includes the following endpoints:

<table data-full-width="false"><thead><tr><th width="271">Endpoint</th><th width="128">Type</th><th>Page</th></tr></thead><tbody><tr><td><code>/createClassification</code></td><td><mark style="color:green;"><code>POST</code></mark></td><td><a data-mention href="/pages/MnkNsxZcjciNT5fYlGNi">/pages/MnkNsxZcjciNT5fYlGNi</a></td></tr><tr><td><code>/viewClassification</code></td><td><mark style="color:green;"><code>POST</code></mark></td><td><a data-mention href="/pages/Be6wluHqDorg6GoCfbTj">/pages/Be6wluHqDorg6GoCfbTj</a></td></tr><tr><td><code>/updateClassification</code></td><td><mark style="color:orange;"><code>PATCH</code></mark></td><td><a data-mention href="/pages/cAfD0cB3isMF3hw2QtFG">/pages/cAfD0cB3isMF3hw2QtFG</a></td></tr><tr><td><code>/deleteExtraction</code></td><td><mark style="color:red;"><code>DELETE</code></mark></td><td><a data-mention href="/pages/Y3K0Cvza86r3bOdI7I1V">/pages/Y3K0Cvza86r3bOdI7I1V</a></td></tr><tr><td><code>/deleteBatch</code></td><td><mark style="color:red;"><code>DELETE</code></mark></td><td><a data-mention href="/pages/serjMgy3W6uoxxDDxY0w">/pages/serjMgy3W6uoxxDDxY0w</a></td></tr><tr><td><code>/deleteFiles</code></td><td><mark style="color:red;"><code>DELETE</code></mark></td><td><a data-mention href="/pages/Hdcz2jJdYtPtLcyKnzhO">/pages/Hdcz2jJdYtPtLcyKnzhO</a></td></tr><tr><td><code>/uploadFiles</code></td><td><mark style="color:green;"><code>POST</code></mark></td><td><a data-mention href="/pages/zn54GLEg771P5xbceEE5">/pages/zn54GLEg771P5xbceEE5</a></td></tr><tr><td><code>/getResults</code></td><td><mark style="color:green;"><code>POST</code></mark></td><td><a data-mention href="/pages/nxZid73okKPSXLrz14f2">/pages/nxZid73okKPSXLrz14f2</a></td></tr></tbody></table>

Each page covers detailed request and response formats.

## 🧩 Document Types & Auto-Extraction

In classification, you define `documentTypes` with:

* A `name` and `description`
* A list of `uniqueWords` to guide classification
* (Optionally) an `extractionId` to enable automatic structured data extraction if a match is found

Learn how to define these properly in the [Document Types](/document-classification-api/classification-details/document-types) page.

## 📘 Navigating the Documentation

Each endpoint in this section includes:

* **Request Parameters:** Fields you can send in the API call
* **Response Objects:** What to expect back
* **Examples:** Realistic payloads to speed up integration

This structure ensures you can build confidently, whether you're uploading 10 documents or 10,000.


# 1. Create classification

<mark style="color:green;">`POST`</mark> `/documentClassification/createClassification`

Initiates a new document classification process. This endpoint allows you to define a classification with a list of possible document types. Once created, you can use the returned `classificationId` to upload documents for type prediction.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="174">Name</th><th width="101">Type</th><th width="104">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>name</code></td><td>string</td><td><code>true</code></td><td>A name for the classification.</td></tr><tr><td><code>description</code></td><td>string</td><td><code>false</code></td><td>A description for the classification.</td></tr><tr><td><code>documentTypes</code></td><td>object</td><td><code>true</code></td><td>An array of objects, each specifying a document type.</td></tr></tbody></table>

## 💡 Need help defining `documentTypes`?

Each item in the `documentTypes` array represents a document category used for classification (e.g., Invoice, Receipt). To learn how to properly define a document type — including required fields, keyword strategy, and optional data extraction linkage — refer to the [Document Types](/document-classification-api/classification-details/document-types) page.

## Body Example

```json
{
  "classificationDetails": {
    "name": "Financial Document Classifier",
    "description": "Classifies uploaded documents into predefined financial document types.",
    "documentTypes": [
      {
        "name": "Invoice",
        "description": "Standard commercial invoice from vendors or suppliers.",
        "uniqueWords": ["invoice number", "bill to", "total amount"],
        "extractionId": "invoiceExtractionId"
      },
      {
        "name": "Purchase Order",
        "description": "Internal or external purchase order documents.",
        "uniqueWords": ["PO number", "item description", "quantity ordered"]
      },
      {
        "name": "Receipt",
        "description": "Retail or online transaction receipts.",
        "uniqueWords": ["receipt", "paid", "transaction id"]
      }
    ]
  }
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

async function createClassification(token, classificationDetails) {
    const url = "https://api.extracta.ai/api/v1/documentClassification/createClassification";

    try {
        const response = await axios.post(url, {
            classificationDetails: classificationDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        return response.data;
    } catch (error) {
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const classificationDetails = {
        name: "Financial Documents Classifier",
        description: "Classifies invoices, receipts, and purchase orders.",
        documentTypes: [
            {
                name: "Invoice",
                description: "Documents with billing and totals for payment.",
                uniqueWords: ["invoice number", "bill to", "total amount"],
                extractionId: "invoiceExtractionId"
            },
            {
                name: "Receipt",
                description: "Confirmation of payment or transaction.",
                uniqueWords: ["receipt", "paid", "transaction id"]
            },
            {
                name: "Purchase Order",
                description: "Authorizes a purchase transaction.",
                uniqueWords: ["PO number", "item", "quantity ordered"]
            }
        ]
    };

    try {
        const response = await createClassification(token, classificationDetails);
        console.log("New Classification Created:", response);
    } catch (error) {
        console.error("Failed to create classification:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_classification(token, classification_details):
    url = "https://api.extracta.ai/api/v1/documentClassification/createClassification"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json={"classificationDetails": classification_details}, headers=headers)
        response.raise_for_status()
        return response.json()
    except requests.RequestException as e:
        print(e)
        return None


if __name__ == "__main__":
    token = "apiKey"
    classification_details = {
        "name": "Financial Documents Classifier",
        "description": "Classifies invoices, receipts, and purchase orders.",
        "documentTypes": [
            {
                "name": "Invoice",
                "description": "Documents with billing and totals for payment.",
                "uniqueWords": ["invoice number", "bill to", "total amount"],
                "extractionId": "invoiceExtractionId"
            },
            {
                "name": "Receipt",
                "description": "Confirmation of payment or transaction.",
                "uniqueWords": ["receipt", "paid", "transaction id"]
            },
            {
                "name": "Purchase Order",
                "description": "Authorizes a purchase transaction.",
                "uniqueWords": ["PO number", "item", "quantity ordered"]
            }
        ]
    }

    response = create_classification(token, classification_details)
    print("New Classification Created:", response)

```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

function createClassification($token, $classificationDetails) {
    $url = 'https://api.extracta.ai/api/v1/documentClassification/createClassification';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = json_encode(['classificationDetails' => $classificationDetails]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_POST, 1);

    try {
        $response = curl_exec($ch);

        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        return $response;
    } catch (Exception $e) {
        return 'Error: ' . $e->getMessage();
    } finally {
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$classificationDetails = [
    "name" => "Financial Documents Classifier",
    "description" => "Classifies invoices, receipts, and purchase orders.",
    "documentTypes" => [
        [
            "name" => "Invoice",
            "description" => "Documents with billing and totals for payment.",
            "uniqueWords" => ["invoice number", "bill to", "total amount"],
            "extractionId" => "invoiceExtractionId"
        ],
        [
            "name" => "Receipt",
            "description" => "Confirmation of payment or transaction.",
            "uniqueWords" => ["receipt", "paid", "transaction id"]
        ],
        [
            "name" => "Purchase Order",
            "description" => "Authorizes a purchase transaction.",
            "uniqueWords" => ["PO number", "item", "quantity ordered"]
        ]
    ]
];

try {
    $response = createClassification($token, $classificationDetails);
    echo $response;
} catch (Exception $e) {
    echo "Failed to create new classification: " . $e->getMessage();
}

?>

```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "created",
    "createdAt": 1712547789609,
    "classificationId": "classificationId"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Name is required"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Error creating classification"
}
```

{% endtab %}
{% endtabs %}


# 2. View classification

<mark style="color:green;">`POST`</mark> `/documentClassification/viewClassification`

This endpoint retrieves the details of a classification process previously defined in the system. By submitting the unique classificationId, you can obtain information such as the classification name, description, configured document types, associated keywords, and any linked extraction templates. This is useful for verifying your classification setup or for debugging and auditing purposes.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="218">Name</th><th width="126">Type</th><th width="115">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>classificationId</code></td><td>string</td><td><code>true</code></td><td>Unique identifier for the classification.</td></tr></tbody></table>

## Body Example

```json
{
    "classificationId": "classificationId"
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

async function viewClassification(token, classificationId) {
    const url = "https://api.extracta.ai/api/v1/documentClassification/viewClassification";

    try {
        const response = await axios.post(url, {
            classificationId: classificationId
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        return response.data;
    } catch (error) {
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const classificationId = 'classificationId';

    try {
        const classificationDetails = await viewClassification(token, classificationId);
        console.log("Classification Details:", classificationDetails);
    } catch (error) {
        console.error("Failed to retrieve classification details:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def view_classification(token, classification_id):   
    url = "https://api.extracta.ai/api/v1/documentClassification/viewClassification"
    
    headers = {
        'Content-Type': 'application/json',
        'Authorization': f'Bearer {token}'
    }
    
    payload = {
        'classificationId': classification_id
    }

    try:
        response = requests.post(url, json=payload, headers=headers)
        response.raise_for_status()
        return response.json()
    except requests.RequestException as e:
        print(f"Failed to retrieve classification details: {e}")
        return None

# Example usage
if __name__ == "__main__":
    token = 'apiKey'
    classification_id = 'classificationId'

    classification_details = view_classification(token, classification_id)
    if classification_details is not None:
        print("Classification Details:", classification_details)

```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

function viewClassification($token, $classificationId) {
    $url = 'https://api.extracta.ai/api/v1/documentClassification/viewClassification';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = json_encode(['classificationId' => $classificationId]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_POST, 1);

    try {
        // Execute cURL session
        $response = curl_exec($ch);

        // Check for cURL errors
        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        // return the response as an associative array
        return json_decode($response, true);
    } catch (Exception $e) {
        return 'Error: ' . $e->getMessage();
    } finally {
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$classificationId = 'classificationId';

try {
    $classificationDetails = viewClassification($token, $classificationId);
    print_r($classificationDetails);
} catch (Exception $e) {
    echo "Failed to retrieve classification details: " . $e->getMessage();
}

?>

```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "success",
    "classificationId": "-OPkce8E1CuEQeHDZetx",
    "classificationDetails": {
        "createdAt": 1746720170530,
        "name": "Financial Document Classifier"
        "description": "Classifies uploaded documents into predefined financial document types.",
        "documentTypes": [
            {
                "description": "Standard commercial invoice from vendors or suppliers.",
                "extractionId": "-OPXR1F82I0cRYJPcHNo",
                "name": "Invoice",
                "uniqueWords": [
                    "invoice number",
                    "bill to",
                    "total amount"
                ]
            },
            {
                "description": "Internal or external purchase order documents.",
                "name": "Purchase Order",
                "uniqueWords": [
                    "PO number",
                    "item description",
                    "quantity ordered"
                ]
            },
            {
                "description": "Retail or online transaction receipts.",
                "name": "Receipt",
                "uniqueWords": [
                    "receipt",
                    "paid",
                    "transaction id"
                ]
            }
        ]
    }
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Classification does not exist"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Error viewing classification"
}
```

{% endtab %}
{% endtabs %}


# 3. Update classification

<mark style="color:orange;">`PATCH`</mark> `/documentClassification/updateClassification`

Updates an existing document classification process by modifying specific parameters within the classification details. This endpoint is useful for adjusting document types, keywords, or metadata without recreating the entire classification setup.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table data-full-width="false"><thead><tr><th width="183">Name</th><th width="125.609375">Type</th><th width="104">Required</th><th width="332.07421875">Description</th></tr></thead><tbody><tr><td><code>classificationId</code></td><td>string</td><td><code>true</code></td><td>Unique identifier for the classification.</td></tr><tr><td><code>name</code></td><td>string</td><td><code>true</code></td><td>A name for the classification.</td></tr><tr><td><code>description</code></td><td>string</td><td><code>true</code></td><td>A description for the classification.</td></tr><tr><td><code>documentTypes</code></td><td>list&#x3C;object></td><td><code>true</code></td><td>An array of objects, each specifying a document type.</td></tr></tbody></table>

## Body Example

```json
{
    "classificationId": "classificationId",
    "classificationDetails": {
        "name": "Financial Document Classifier - updated",
        "description": "Classifies uploaded documents into predefined financial document types. - updated",
        "documentTypes": [
            {
                "name": "Invoice",
                "description": "Standard commercial invoice from vendors or suppliers.",
                "uniqueWords": [
                    "invoice number",
                    "bill to",
                    "total amount"
                ],
                "extractionId": "-OPXR1F82I0cRYJPcHNo"
            }
        ]
    }
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

async function updateClassification(token, classificationId, classificationDetails) {
    const url = "https://api.extracta.ai/api/v1/documentClassification/updateClassification";

    try {
        const response = await axios.patch(url, {
            classificationId,
            classificationDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        return response.data;
    } catch (error) {
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const classificationId = 'classificationId';
    const classificationDetails = {
        "name": "Financial Document Classifier - updated",
        "description": "Classifies uploaded documents into predefined financial document types. - updated",
        "documentTypes": [
            {
                "name": "Invoice",
                "description": "Standard commercial invoice from vendors or suppliers.",
                "uniqueWords": [
                    "invoice number",
                    "bill to",
                    "total amount"
                ],
                "extractionId": "-OPXR1F82I0cRYJPcHNo"
            }
        ]
    };

    try {
        const response = await updateClassification(token, classificationId, classificationDetails);
        console.log("Classification Updated:", response);
    } catch (error) {
        console.error("Failed to update classification:", error);
    }
}

main();

```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def update_classification(token, classification_id, classification_details):
    url = "https://api.extracta.ai/api/v1/documentClassification/updateClassification"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}
    payload = {
        "classificationId": classification_id,
        "classificationDetails": classification_details
    }

    try:
        response = requests.patch(url, json=payload, headers=headers)
        response.raise_for_status()
        return response.json()
    except requests.RequestException as e:
        print(f"Failed to update classification: {e}")
        return None

# Example usage
if __name__ == "__main__":
    token = "apiKey"
    classification_id = "classificationId"
    classification_details = {
        "name": "Financial Document Classifier - updated",
        "description": "Classifies uploaded documents into predefined financial document types. - updated",
        "documentTypes": [
            {
                "name": "Invoice",
                "description": "Standard commercial invoice from vendors or suppliers.",
                "uniqueWords": [
                    "invoice number",
                    "bill to",
                    "total amount"
                ],
                "extractionId": "-OPXR1F82I0cRYJPcHNo"
            }
        ]
    }

    response = update_classification(token, classification_id, classification_details)
    print("Classification Updated:", response)

```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

function updateClassification($token, $classificationId, $classificationDetails) {
    $url = 'https://api.extracta.ai/api/v1/documentClassification/updateClassification';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload with classificationId and classificationDetails
    $payload = json_encode([
        'classificationId' => $classificationId,
        'classificationDetails' => $classificationDetails
    ]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_CUSTOMREQUEST, 'PATCH');
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

    try {
        // Execute cURL session
        $response = curl_exec($ch);

        // Check for cURL errors
        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        return $response;
    } catch (Exception $e) {
        return 'Error: ' . $e->getMessage();
    } finally {
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$classificationId = 'classificationId';
$classificationDetails = [
    "name" => "Updated Financial Classifier",
    "description" => "Updated description for the financial document classifier.",
    "documentTypes" => [
        [
            "name" => "Invoice",
            "description" => "Updated invoice description.",
            "uniqueWords" => ["invoice number", "billing address"],
            "extractionId" => "updatedExtract123"
        ],
        [
            "name" => "Receipt",
            "description" => "Retail receipts for in-store purchases.",
            "uniqueWords" => ["store id", "cashier", "receipt total"]
        ]
    ]
];

try {
    $response = updateClassification($token, $classificationId, $classificationDetails);
    echo $response;
} catch (Exception $e) {
    echo "Failed to update classification: " . $e->getMessage();
}

?>

```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "updated",
    "updatedAt": 1746720927500,
    "classificationId": "-OPkce8E1CuEQeHDZetx"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Classification does not exist",
    "classificationId": "-OPkce8E1CuEQeHDZetxs"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Error updating classification"
}
```

{% endtab %}
{% endtabs %}


# 4. Delete data


# 4.1 Delete classification

<mark style="color:red;">`DELETE`</mark> `/documentClassification/deleteClassification`

This endpoint permanently deletes an entire document classification process, including all associated batches, results, and uploaded files.

You must provide only the `classificationId` in the request body. This action is irreversible and removes all data linked to the specified classification.

Use this endpoint when you intend to fully decommission a classification and all its related content. Separate endpoints are available for deleting individual batches or files.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="183">Name</th><th width="135">Type</th><th width="164">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>classificationId</code></td><td>string</td><td><code>true</code></td><td>The classification id.</td></tr></tbody></table>

## Body Example

```json
{
    "classificationId": "classificationId"
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

async function deleteClassification(token, classificationId) {
    const url = "https://api.extracta.ai/api/v1/documentClassification/deleteClassification";

    try {
        const response = await axios.delete(url, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            },
            data: {
                classificationId: classificationId
            }
        });

        return response.data;
    } catch (error) {
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const classificationId = 'classificationId';

    try {
        const response = await deleteClassification(token, classificationId);
        console.log("Classification Deleted:", response);
    } catch (error) {
        console.error("Failed to delete classification:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def delete_classification(token, classification_id):
    url = "https://api.extracta.ai/api/v1/documentClassification/deleteClassification"
    headers = {
        "Content-Type": "application/json",
        "Authorization": f"Bearer {token}"
    }
    payload = {
        "classificationId": classification_id
    }

    try:
        response = requests.delete(url, json=payload, headers=headers)
        response.raise_for_status()
        return response.json()
    except requests.RequestException as e:
        print(f"Failed to delete classification: {e}")
        return None

# Example usage
if __name__ == "__main__":
    token = "apiKey"
    classification_id = "classificationId"

    response = delete_classification(token, classification_id)
    if response:
        print("Classification Deleted:", response)
    else:
        print("Failed to delete classification.")
```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

function deleteClassification($token, $classificationId) {
    $url = 'https://api.extracta.ai/api/v1/documentClassification/deleteClassification';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = json_encode(['classificationId' => $classificationId]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_CUSTOMREQUEST, 'DELETE');
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

    try {
        $response = curl_exec($ch);

        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        return $response;
    } catch (Exception $e) {
        return 'Error: ' . $e->getMessage();
    } finally {
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$classificationId = 'classificationId';

try {
    $response = deleteClassification($token, $classificationId);
    echo "Classification Deleted: " . $response;
} catch (Exception $e) {
    echo "Failed to delete classification: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "success",
    "message": "Classification deleted"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "success",
    "message": "Classification deleted"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Internal server error"
}
```

{% endtab %}
{% endtabs %}


# 4.2 Delete batch

<mark style="color:red;">`DELETE`</mark> `/documentClassification/deleteBatch`

This endpoint permanently deletes a specific batch from a document classification process.

You must provide both the `classificationId` and the `batchId` in the request body. This action is irreversible and will remove the selected batch along with all results and files associated with it.

Use this endpoint when you need to clean up or remove a specific batch without affecting the overall classification. Separate endpoints are available for deleting entire classifications or individual files.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="183">Name</th><th width="135">Type</th><th width="164">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>classificationId</code></td><td>string</td><td><code>true</code></td><td>The classification id.</td></tr><tr><td><code>batchId</code></td><td>string</td><td><code>true</code></td><td>The batch id.</td></tr></tbody></table>

## Body Example

```json
{
    "classificationId": "classificationId",
    "batchId": "batchId"
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

async function deleteBatch(token, classificationId, batchId) {
    const url = "https://api.extracta.ai/api/v1/documentClassification/deleteBatch";

    try {
        const response = await axios.delete(url, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            },
            data: {
                classificationId: classificationId,
                batchId: batchId
            }
        });

        return response.data;
    } catch (error) {
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const classificationId = 'classificationId';
    const batchId = 'batchId';

    try {
        const response = await deleteBatch(token, classificationId, batchId);
        console.log("Batch Deleted:", response);
    } catch (error) {
        console.error("Failed to delete batch:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def delete_batch(token, classification_id, batch_id):
    url = "https://api.extracta.ai/api/v1/documentClassification/deleteBatch"
    headers = {
        "Content-Type": "application/json",
        "Authorization": f"Bearer {token}"
    }
    payload = {
        "classificationId": classification_id,
        "batchId": batch_id
    }

    try:
        response = requests.delete(url, json=payload, headers=headers)
        response.raise_for_status()
        return response.json()
    except requests.RequestException as e:
        print(f"Failed to delete classification: {e}")
        return None

# Example usage
if __name__ == "__main__":
    token = "apiKey"
    classification_id = "classificationId"
    batch_id = "batchId"

    response = delete_batch(token, classification_id, batch_id)
    if response:
        print("Batch Deleted:", response)
    else:
        print("Failed to delete batch.")
```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

function deleteBatch($token, $classificationId, $batchId) {
    $url = 'https://api.extracta.ai/api/v1/documentClassification/deleteBatch';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = json_encode([
        'classificationId' => $classificationId,
        'batchId' => $batchId
    ]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_CUSTOMREQUEST, 'DELETE');
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

    try {
        $response = curl_exec($ch);

        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        return $response;
    } catch (Exception $e) {
        return 'Error: ' . $e->getMessage();
    } finally {
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$classificationId = 'classificationId';
$batchId = 'batchId';

try {
    $response = deleteBatch($token, $classificationId, $batchId);
    echo "Batch Deleted: " . $response;
} catch (Exception $e) {
    echo "Failed to delete batch: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "success",
    "message": "Batch deleted"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "success",
    "message": "Batch deleted"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Internal server error"
}
```

{% endtab %}
{% endtabs %}


# 4.3 Delete files

<mark style="color:red;">`DELETE`</mark> `/documentClassification/deleteFiles`

This endpoint permanently deletes one or more specific files from a batch within a document classification process.

You must provide the `classificationId`, `batchId`, and a list of `fileIds` in the request body. This action is irreversible and will remove the specified files and their associated results, without affecting the rest of the batch or classification.

Use this endpoint to precisely remove files without deleting the entire batch or classification. Separate endpoints exist for deleting entire classifications or batches.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="183">Name</th><th width="135">Type</th><th width="164">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>classificationId</code></td><td>string</td><td><code>true</code></td><td>The classification id.</td></tr><tr><td><code>batchId</code></td><td>string</td><td><code>true</code></td><td>The batch id.</td></tr><tr><td><code>fileIds</code></td><td>list&#x3C;string></td><td><code>true</code></td><td>A list of specific file ids.</td></tr></tbody></table>

## Body Example

```json
{
    "classificationId": "classificationId",
    "batchId": "batchId",
    "fileIds": ["fileId1", "fileId2", "fileId3"]
}
```

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

async function deleteFiles(token, classificationId, batchId, fileIds) {
    const url = "https://api.extracta.ai/api/v1/documentClassification/deleteFiles";

    try {
        const response = await axios.delete(url, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            },
            data: {
                classificationId: classificationId,
                batchId: batchId,
                fileIds: fileIds
            }
        });

        return response.data;
    } catch (error) {
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const classificationId = 'classificationId';
    const batchId = 'batchId';
    const fileIds = ["fileId1", "fileId2", "fileId3"];

    try {
        const response = await deleteFiles(token, classificationId, batchId, fileIds);
        console.log("Files Deleted:", response);
    } catch (error) {
        console.error("Failed to delete files:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def delete_files(token, classification_id, batch_id, file_ids):
    url = "https://api.extracta.ai/api/v1/documentClassification/deleteFiles"
    headers = {
        "Content-Type": "application/json",
        "Authorization": f"Bearer {token}"
    }
    payload = {
        "classificationId": classification_id,
        "batchId": batch_id,
        "fileIds": file_ids
    }

    try:
        response = requests.delete(url, json=payload, headers=headers)
        response.raise_for_status()
        return response.json()
    except requests.RequestException as e:
        print(f"Failed to delete classification: {e}")
        return None

# Example usage
if __name__ == "__main__":
    token = "apiKey"
    classification_id = "classificationId"
    batch_id = "batchId"
    file_ids = ["fileId1", "fileId2", "fileId3"]

    response = delete_batch(token, classification_id, batch_id, file_ids)
    if response:
        print("Files Deleted:", response)
    else:
        print("Failed to delete files.")
```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

function deleteFiles($token, $classificationId, $batchId, $fileIds) {
    $url = 'https://api.extracta.ai/api/v1/documentClassification/deleteFiles';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = json_encode([
        'classificationId' => $classificationId,
        'batchId' => $batchId,
        'fileIds' => $fileIds
    ]);

    // Set cURL options
    curl_setopt($ch, CURLOPT_CUSTOMREQUEST, 'DELETE');
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

    try {
        $response = curl_exec($ch);

        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        return $response;
    } catch (Exception $e) {
        return 'Error: ' . $e->getMessage();
    } finally {
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$classificationId = 'classificationId';
$batchId = 'batchId';
$fileIds = ['fileId1', 'fileId2', 'fileId3'];

try {
    $response = deleteFiles($token, $classificationId, $batchId, $fileIds);
    echo "Files Deleted: " . $response;
} catch (Exception $e) {
    echo "Failed to delete files: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "success",
    "message": "Files deleted"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "success",
    "message": "Files deleted"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Internal server error"
}
```

{% endtab %}
{% endtabs %}


# 5. Upload Files

<mark style="color:green;">`POST`</mark> `/documentClassification/deleteFiles`

This endpoint enables users to upload files to a specific document classification process.

If a `batchId` is provided in the request, the uploaded files will be added to that existing batch. If `batchId` is omitted, a new batch will be automatically created for the files under the specified `classificationId`.

Files must be uploaded using the `multipart/form-data` content type, which is required for binary file uploads such as PDFs, images, or scanned documents. Ensure the `classificationId` is valid, and that any provided `batchId` already exists within the classification context.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value                 |
| ------------- | --------------------- |
| Content-Type  | `multipart/form-data` |
| Authorization | `Bearer <token>`      |

## Body

<table><thead><tr><th width="251">Name</th><th width="119">Type</th><th width="115">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>classificationId</code></td><td>string</td><td><code>true</code></td><td>Unique identifier for the extraction.</td></tr><tr><td><code>batchId</code></td><td>string</td><td><code>false</code></td><td>The ID of the batch to add files to.</td></tr><tr><td><code>files</code></td><td>multipart</td><td><code>true</code></td><td><a data-mention href="/pages/GEp82eoMVXcJO7dA0Z4j">/pages/GEp82eoMVXcJO7dA0Z4j</a></td></tr></tbody></table>

For a seamless classification process, please ensure your documents are in one of our supported formats. Check our Supported File Types page for a list of all formats we currently accept and additional details to prepare your files accordingly.

{% content-ref url="/pages/GEp82eoMVXcJO7dA0Z4j" %}
[Broken mention](broken://pages/GEp82eoMVXcJO7dA0Z4j)
{% endcontent-ref %}

## Code Example

**Note for PHP Users:** Currently, the `/uploadFiles` endpoint supports uploading only one file per request. Please ensure you submit individual requests for each file you need to upload.

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const fs = require('fs');
const axios = require('axios');
const FormData = require('form-data');

async function uploadFilesToClassification(token, classificationId, files, batchId = null) {
    const url = "https://api.extracta.ai/api/v1/documentClassification/uploadFiles";
    let formData = new FormData();

    formData.append('classificationId', classificationId);
    if (batchId) {
        formData.append('batchId', batchId);
    }

    files.forEach(file => {
        formData.append('files', fs.createReadStream(file));
    });

    try {
        const response = await axios.post(url, formData, {
            headers: {
                ...formData.getHeaders(),
                'Authorization': `Bearer ${token}`
            }
        });

        return response.data;
    } catch (error) {
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const classificationId = 'classificationId';
    const files = ['test_1.pdf', 'test_2.jpg'];
    const batchId = null; // or specify an existing batchId if needed

    try {
        const response = await uploadFilesToClassification(token, classificationId, files, batchId);
        console.log("Upload Response:", response);
    } catch (error) {
        console.error("Failed to upload files:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests
import mimetypes

def upload_files_to_classification(token, classification_id, files, batch_id=None):
    url = "https://api.extracta.ai/api/v1/documentClassification/uploadFiles"
    headers = {"Authorization": f"Bearer {token}"}

    # Prepare the files for uploading
    file_streams = [
        (
            "files",
            (
                file,
                open(file, "rb"),
                mimetypes.guess_type(file)[0] or "application/octet-stream",
            ),
        )
        for file in files
    ]
    
    payload = {"classificationId": classification_id}
    if batch_id is not None:
        payload["batchId"] = batch_id

    try:
        response = requests.post(url, files=file_streams, data=payload, headers=headers)
        response.raise_for_status()
        return response.json()
    except requests.HTTPError as e:
        if response.status_code >= 400:
            try:
                print("Server returned an error:", response.json())
            except Exception:
                print("Server returned an error:", response.text)
        else:
            print(f"HTTP error occurred: {e}")
    except requests.RequestException as e:
        print(f"Failed to upload files: {e}")
    except Exception as e:
        print(f"An unexpected error occurred: {e}")
    return None

# Example usage
if __name__ == "__main__":
    token = 'apiKey'
    classification_id = 'classificationId'
    files = ['test_1.pdf', 'test_2.jpg']
    batch_id = None  # or specify an existing batch ID if needed

    response = upload_files_to_classification(token, classification_id, files, batch_id)
    if response:
        print("Upload response:", response)
    else:
        print("Upload failed.")
```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

function uploadFileToClassification($token, $classificationId, $filePath, $batchId = null) {
    $url = 'https://api.extracta.ai/api/v1/documentClassification/uploadFiles';

    // Initialize cURL session
    $ch = curl_init($url);

    // Prepare the payload
    $payload = [
        'classificationId' => $classificationId,
        'files' => new CURLFile($filePath, mime_content_type($filePath), basename($filePath))
    ];

    if ($batchId !== null) {
        $payload['batchId'] = $batchId;
    }

    // Set cURL options
    curl_setopt($ch, CURLOPT_POST, 1);
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Authorization: Bearer ' . $token,
        'Content-Type: multipart/form-data'
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);

    try {
        $response = curl_exec($ch);

        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        return $response;
    } catch (Exception $e) {
        return 'Error: ' . $e->getMessage();
    } finally {
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$classificationId = 'yourClassificationIdHere';
$filePath = './test_1.pdf';
$batchId = null; // or specify an existing batch ID

try {
    $response = uploadFileToClassification($token, $classificationId, $filePath, $batchId);
    echo $response;
} catch (Exception $e) {
    echo "Failed to upload file: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

{% tabs %}
{% tab title="200" %}

```json
{
    "status": "uploaded",
    "classificationId": "classificationId",
    "batchId": "batchId",
    "files": [
        {
            "fileId": "fileId",
            "fileName": "file.pdf",
            "numberOfPages": 1,
            "url": "fileUrl"
        }
    ]
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Classification not found"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
    "status": "error",
    "message": "Internal server error"
}
```

{% endtab %}
{% endtabs %}


# 6. Get results

<mark style="color:green;">`POST`</mark> `/documentClassification/getResults`

This endpoint retrieves the classification results for a specific batch of documents.

By providing the `classificationId` and `batchId`, you can obtain the predicted document types for each file in the batch,.

Optionally, including a `fileId` will return results for that specific file only. This is useful for retrieving targeted results without querying the entire batch.

## Server URL

```
https://api.extracta.ai/api/v1
```

## Headers

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

## Body

<table><thead><tr><th width="242">Name</th><th width="99">Type</th><th width="126">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>classificationId</code></td><td>string</td><td><code>true</code></td><td>The ID of the classification.</td></tr><tr><td><code>batchId</code></td><td>string</td><td><code>true</code></td><td>The ID of the batch.</td></tr><tr><td><code>fileId</code></td><td>string</td><td><code>false</code></td><td>The ID of the file.</td></tr></tbody></table>

## Body Example

```json
{
    "classificationId": "classificationId",
    "batchId": "batchId",
    "fileId": "fileId" // optional
}
```

## ⚠️ Important

To avoid rate-limiting, please ensure a delay of 2 seconds between consecutive requests to this endpoint.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

async function getClassificationResults(token, classificationId, batchId, fileId) {
    const url = "https://api.extracta.ai/api/v1/documentClassification/getResults";

    try {
        const payload = {
            classificationId,
            batchId
        };

        if (fileId) {
            payload.fileId = fileId;
        }

        const response = await axios.post(url, payload, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        return response.data;
    } catch (error) {
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const classificationId = 'classificationId';
    const batchId = 'batchId';
    const fileId = null; // or provide a specific file ID

    try {
        const results = await getClassificationResults(token, classificationId, batchId, fileId);
        console.log("Classification Results:", results);
    } catch (error) {
        console.error("Failed to retrieve classification results:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests

def get_classification_results(token, classification_id, batch_id, file_id=None):
    url = "https://api.extracta.ai/api/v1/documentClassification/getResults"
    
    headers = {
        'Content-Type': 'application/json',
        'Authorization': f'Bearer {token}'
    }
    
    payload = {
        'classificationId': classification_id,
        'batchId': batch_id
    }

    if file_id:
        payload['fileId'] = file_id

    response = requests.post(url, json=payload, headers=headers)
    
    if response.status_code == 200:
        return response.json()
    else:
        response.raise_for_status()

# Example usage
if __name__ == "__main__":
    token = 'apiKey'
    classification_id = 'classificationId'
    batch_id = 'batchId'
    file_id = None  # Or use a specific file ID like 'file123'

    try:
        results = get_classification_results(token, classification_id, batch_id, file_id)
        print("Classification Results:", results)
    except Exception as e:
        print("Failed to retrieve classification results:", e)

```

{% endtab %}

{% tab title="PHP" %}

```php
<?php

function getClassificationResults($token, $classificationId, $batchId, $fileId = null) {
    $url = 'https://api.extracta.ai/api/v1/documentClassification/getResults';

    // Initialize cURL session
    $ch = curl_init($url);
    
    // Prepare the payload
    $payload = [
        'classificationId' => $classificationId,
        'batchId' => $batchId
    ];

    if ($fileId !== null) {
        $payload['fileId'] = $fileId;
    }

    $payload = json_encode($payload);

    // Set cURL options
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $token,
    ]);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_POST, 1);

    try {
        $response = curl_exec($ch);

        if (curl_errno($ch)) {
            throw new Exception('Curl error: ' . curl_error($ch));
        }

        return $response;
    } catch (Exception $e) {
        return 'Error: ' . $e->getMessage();
    } finally {
        curl_close($ch);
    }
}

// Example usage
$token = 'apiKey';
$classificationId = 'classificationId';
$batchId = 'batchId';
$fileId = null; // or 'yourFileId' if you want to filter by file

try {
    $results = getClassificationResults($token, $classificationId, $batchId, $fileId);
    echo $results;
} catch (Exception $e) {
    echo "Failed to retrieve classification results: " . $e->getMessage();
}

?>
```

{% endtab %}
{% endtabs %}

## Responses

In case a file is linked to an extraction template, the response will also include an **`extraction`** object with:

* `extractionId` → the ID of the created extraction job
* `batchId` → the batch in which the file was processed
* `fileId` → the ID of the file that was uploaded and processed in the **data extraction module**

{% tabs %}
{% tab title="200 for all files" %}

```json
{
    "status": "success",
    "classificationId": "classificationId",
    "batchId": "batchId",
    "files": [
        {
            "fileId": "fileId1",
            "fileName": "File 1.pdf",
            "status": "processed",
            "result": {
                "confidence": 0.9,
                "documentType": "invoice"
            },
            "url": "fileUrl"
        },
        {
            "fileId": "fileId1",
            "fileName": "File 2.pdf",
            "status": "processed",
            "result": {
                "confidence": 0.95,
                "documentType": "resume"
            },
            "extraction": {
                "extractionId": "extractionId",
                "batchId": "batchId",
                "fileId": "fileId"
            },
            "url": "fileUrl"
        },
        {
            "fileId": "fileId2",
            "fileName": "File 3.pdf",
            "status": "processed",
            "result": {
                "confidence": 0.95,
                "documentType": "other"
            },
            "url": "fileUrl"
        },
        ...
    ]
}
```

{% endtab %}

{% tab title="200 for a single file" %}

```json
{
    "status": "success",
    "classificationId": "classificationId",
    "batchId": "batchId",
    "fileId": "fileId",
    "files": [
        {
            "fileId": "fileId",
            "fileName": "File 1.pdf",
            "status": "processed",
            "result": {
                "confidence": 0.9,
                "documentType": "invoice"
            },
            "url": "fileUrl"
        }
    ]
}
```

{% endtab %}

{% tab title="200 status waiting" %}

```json
{
    "status": "waiting",
    "classificationId": "-OPXQuIrkHuA2QZkY-44",
    "batchId": "8pMpPs6SMPBpcnnznoaRz8bsj"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
    "status": "error",
    "message": "Internal server error"
}
```

{% endtab %}
{% endtabs %}


# Classification Details


# Document Types

In Extracta's document classification system, a `Document Type` defines a category of documents you expect to classify—such as invoices, receipts, contracts, or purchase orders. Each document type helps the system determine how to interpret and route the documents you upload.

## 🧩 Whats a Document Type?

A documentType is a JSON object with the following fields:

<table><thead><tr><th width="179.61328125">Field</th><th width="161.2734375">Type</th><th width="104.8828125">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>name</code></td><td><code>string</code></td><td>✅</td><td>A clear, human-readable label for the document type (e.g., <code>"Invoice"</code>).</td></tr><tr><td><code>description</code></td><td><code>string</code></td><td>✅</td><td>A short explanation of the type of documents this represents.</td></tr><tr><td><code>uniqueWords</code></td><td><code>list&#x3C;string></code></td><td>✅</td><td>A list of key terms or phrases likely to appear in this document type.</td></tr><tr><td><code>extractionId</code></td><td><code>string</code></td><td>❌</td><td>ID of a pre-configured extraction template to auto-extract data.</td></tr></tbody></table>

## 📌 Why Unique Words Matter?

The `uniqueWords` array is critical to helping our classification model distinguish between document types. These are **keywords or phrases** commonly found in that document type. Think of them like clues for classification.

```json
"uniqueWords": [
    "invoice number", 
    "bill to", 
    "total amount"
]
```

This helps the classifier recognize invoices based on visible cues.

## ⚙️ Linking to Extraction Templates

You can optionally link a `documentType` to an existing **extraction template** using `extractionId`. This allows you to:

1. Upload a batch of documents for classification.
2. Have each document automatically classified (e.g., as an "Invoice").
3. Automatically trigger **data extraction** for that document using the associated template.

This creates a powerful end-to-end flow: **classification ➝ structured data extraction**.

## 🧾 Example Definition

Here’s how a `documentTypes` array might look:

```json
"documentTypes": [
  {
    "name": "Invoice",
    "description": "Standard commercial invoice from vendors or suppliers.",
    "uniqueWords": ["invoice number", "bill to", "total amount"],
    "extractionId": "invoiceExtractionId"
  },
  {
    "name": "Purchase Order",
    "description": "Internal or external purchase order documents.",
    "uniqueWords": ["PO number", "item description", "quantity ordered"]
  },
  {
    "name": "Receipt",
    "description": "Retail or online transaction receipts.",
    "uniqueWords": ["receipt", "paid", "transaction id"]
  }
]
```


# Postman Integration

<div align="left"><figure><img src="/files/y2yBc3LbuDb3UnlwB0pz" alt=""><figcaption></figcaption></figure></div>

To help you get started quickly with the Extracta.ai API, we have created a comprehensive Postman collection. This collection includes all the API endpoints, complete with example requests and responses, making it easier for you to test and integrate with our API.

### How to Use the Postman Collection

* **Import the Collection into your Postman account:** <https://www.postman.com/extracta/workspace/extracta-ai-workspace/collection/35309516-6de7ea3c-857e-486e-b721-7488189881bf?action=share&creator=35309516>
* **Set Up Your Environment:**
  * Configure your variables in **Collection -> Variables** to include your **API key**.
  * How to make an API Key: [Authentication](/api-reference/authentication)

<figure><img src="/files/AqwQvYSrAWlm9QcRO2YW" alt=""><figcaption></figcaption></figure>

* **Explore the API Endpoints:** You can easily navigate through the endpoints and see example requests and responses.

### Import directly the collection by JSON File.

* **Download JSON File with API Endpoints**: <https://drive.google.com/file/d/1PLK1a2-LPr30XBfEdJYaCNmiFx34J-PU/view?usp=sharing>
* **Import the file in Postman as in the screenshot below:**

<figure><img src="/files/oaTqdKlJ85dMAK9RY1d4" alt=""><figcaption></figcaption></figure>

### Benefits of Using the Postman Collection

* **Quick Start:** Instantly access all available API endpoints with pre-configured requests.
* **Ease of Use:** Easily test different endpoints and see example responses without writing any code.
* **Consistency:** Ensure that you are using the correct request formats and parameters as per our latest API specifications.

### Additional Resources

* **API Endpoints Page:** [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction)
* **Support:** If you encounter any issues or have questions, please visit [Contact Us](/contact/contact-us) page.

We hope this Postman collection makes your development process smoother and more efficient. Happy coding!


# Custom Document

Extracta LABS empowers businesses to automate the extraction of structured and unstructured data from a wide array of document types. Leveraging AI-powered OCR technology, our API facilitates the transformation of scanned documents in PDF, JPG, and PNG formats into actionable, intelligent data. This capability is crucial for businesses looking to digitalize their operations through advanced data analysis, mining, and Named Entity Recognition (NER). Whether you're processing individual files or handling documents in batches, Extracta.ai offers a streamlined solution for your automation and information retrieval needs.

## Why Choose Extracta LABS for Custom Document Parsing?

* **Versatile Data Extraction**: Tailor the API to meet your specific document parsing requirements, whether your documents are structured or unstructured.
* **Advanced OCR Technology**: Convert scanned documents into digital data with high accuracy, making your data easily accessible and actionable.
* **Intelligent Data Processing**: Utilize IDP and NLP technologies for deeper insights, enabling effective data analysis and decision-making processes.
* **Batch File Processing**: Efficiently process large volumes of documents, saving time and resources while enhancing operational efficiency.

## Getting Started

To harness the power of custom document parsing, you'll start by crafting a `POST` request to our `/createExtraction` endpoint. The key to leveraging this functionality lies in specifying the custom fields relevant to your documents. Here’s how you can structure your request body for a custom document type.

## Body Example for Custom Document Parsing

<details>

<summary>JSON Body</summary>

```json
{
    "extractionDetails": {
        "name": "Custom document - Extraction",
        "language": "English",
        "fields": [
            {
                "key": "name",
                "description": "the name of the person in the CV",
                "example": "Johan Smith"
            },
            {
                "key": "email",
                "description": "the email of the person in the CV",
                "example": "johan@gmail.com"
            },
            {
                "key": "phone",
                "description": "the phone number of the person",
                "example": "123 333 4445"
            },
            {
                "key": "address",
                "description": "the compelte address of the person",
                "example": "1234 Main St, New York, NY 10001"
            },
            {
                "key": "soft_skills",
                "description": "the soft skills of the person",
                "example": ""
            },
            {
                "key": "hard_skills",
                "description": "the hard skills of the person",
                "example": ""
            },
            {
                "key": "last_job",
                "description": "the last job of the person",
                "example": "Software Engineer"
            },
            {
                "key": "years_of_experience",
                "description": "the years of experience of last job",
                "example": "5"
            }
        ]
    }
}
```

</details>

## Customizing Your Request

1. **Define Your Fields**: These fields should reflect the unique aspects of your custom document type.
2. **Prepare the API Call**: Incorporate the JSON template into your `POST` request body to `/createExtraction`. Ensure your API key is included in the header as a Bearer token for authentication.
3. **Process the Extracted Data**: After submitting your request, the API will analyze your document and return structured data based on the custom fields you defined, ready for integration into your systems.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {}; // the json body from the example

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {} # the json body from the example

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}
{% endtabs %}

## Conclusion

With Extracta LABS, the complexity of parsing custom document types is simplified, enabling businesses to focus on deriving value from their data rather than on the intricacies of data extraction. Our customizable approach ensures that you can adapt the API to fit the unique needs of your documents, facilitating efficient data handling and analysis.

For detailed guidance on API integration and making the most of Extracta LABS’s capabilities, please explore our [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction) and [1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction) pages.


# Resume / CV

Unlock the full potential of resume data with Extracta LABS’s advanced document parsing capabilities. Our API simplifies the process of extracting valuable information from resumes, making it easier for HR professionals, recruiters, and job portals to automate candidate data extraction, enhancing the recruitment process. This page provides you with a ready-to-use request body for creating resume extractions, streamlining your integration with our service.

## Why Use Extracta LABS for Resume Parsing?

* **Automated Data Extraction**: Reduce manual data entry and speed up the candidate screening process.
* **High Accuracy**: Leverage Extracta LABS's advanced algorithms to ensure precise extraction of candidate information.
* **Customizable Extraction**: Tailor the data extraction to meet your specific recruitment needs.

## Getting Started

To begin parsing resumes with Extracta LABS, you'll need to make a `POST` request to our `/createExtraction` endpoint. Below is a template for the request body specifically designed for resume parsing. This template includes predefined keys that correspond to common data points found in resumes, such as personal information, education, work experience, and skills.

## Body Example for Resume / CV

<details>

<summary>JSON Body</summary>

```json
{
    "extractionDetails": {
        "name": "Resume - Extraction",
        "language": "English",
        "fields": [
            {
                "key": "personal_info",
                "description": "personal information of the person",
                "type": "object",
                "properties": [
                    {
                        "key": "name",
                        "description": "name of the person",
                        "example": "Alex Smith",
                        "type": "string"
                    },
                    {
                        "key": "email",
                        "description": "email of the person",
                        "example": "alex.smith@gmail.com",
                        "type": "string"
                    },
                    {
                        "key": "phone",
                        "description": "phone of the person",
                        "example": "0712 123 123",
                        "type": "string"
                    },
                    {
                        "key": "address",
                        "description": "address of the person",
                        "example": "Bucharest, Romania",
                        "type": "string"
                    }
                ]
            },
            {
                "key": "work_experience",
                "description": "work experience of the person",
                "type": "array",
                "items": {
                    "type": "object",
                    "properties": [
                        {
                            "key": "title",
                            "description": "title of the job",
                            "example": "Software Engineer",
                            "type": "string"
                        },
                        {
                            "key": "start_date",
                            "description": "start date of the job",
                            "example": "2022",
                            "type": "string"
                        },
                        {
                            "key": "end_date",
                            "description": "end date of the job",
                            "example": "2023",
                            "type": "string"
                        },
                        {
                            "key": "company",
                            "description": "company of the job",
                            "example": "Fastapp Development",
                            "type": "string"
                        },
                        {
                            "key": "location",
                            "description": "location of the job",
                            "example": "Bucharest, Romania",
                            "type": "string"
                        },
                        {
                            "key": "description",
                            "description": "description of the job",
                            "example": "Designing and implementing server-side logic to ensure high performance and responsiveness of applications.",
                            "type": "string"
                        }
                    ]
                }
            },
            {
                "key": "education",
                "description": "school education of the person",
                "type": "array",
                "items": {
                    "type": "object",
                    "properties": [
                        {
                            "key": "title",
                            "description": "title of the education",
                            "example": "Master of Science in Computer Science",
                            "type": "string"
                        },
                        {
                            "key": "start_date",
                            "description": "start date of the education",
                            "example": "2022",
                            "type": "string"
                        },
                        {
                            "key": "end_date",
                            "description": "end date of the education",
                            "example": "2023",
                            "type": "string"
                        },
                        {
                            "key": "institute",
                            "description": "institute of the education",
                            "example": "Bucharest Academy of Economic Studies",
                            "type": "string"
                        },
                        {
                            "key": "location",
                            "description": "location of the education",
                            "example": "Bucharest, Romania",
                            "type": "string"
                        },
                        {
                            "key": "description",
                            "description": "description of the education",
                            "example": "Advanced academic degree focusing on developing a deep understanding of theoretical foundations and practical applications of computer technology.",
                            "type": "string"
                        }
                    ]
                }
            },
            {
                "key": "languages",
                "description": "languages spoken by the person",
                "type": "array",
                "items": {
                    "type": "string",
                    "example": "English"
                }
            },
            {
                "key": "skills",
                "description": "skills of the person",
                "type": "array",
                "items": {
                    "type": "string",
                    "example": "NodeJS"
                }
            },
            {
                "key": "certificates",
                "description": "certificates of the person",
                "type": "array",
                "items": {
                    "type": "string",
                    "example": "AWS Certified Developer - Associate"
                }
            }
        ]
    }
}
```

</details>

## How to Use This Template

1. **Create an API Key**: Ensure you have an API key by signing up at [https://app.extracta.ai](https://app.extracta.ai/) and navigating to the `/api` page to generate your key.
2. **Make the API Call**: Use the provided JSON template in your `POST` request to `/createExtraction`. Don’t forget to include your API key in the request header as a Bearer token for authentication.
3. **Receive and Process Data**: After submitting your request, the API will return structured data extracted from the resume, which you can then integrate into your application or workflow.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {}; // the json body from the example

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {} # the json body from the example

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}
{% endtabs %}

## Conclusion

Extracta LABS streamlines the extraction of critical information from resumes, offering a powerful tool for anyone looking to automate and improve the efficiency of their recruitment process. By utilizing our API, you can focus on what truly matters - finding the right candidates.

For further details on integrating our API and making the most out of its capabilities, please refer to our [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction) and [1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction) pages.


# Contract

Maximize the efficiency of contract management with Extracta LABS’s cutting-edge document parsing technology. Our API is designed to automate the extraction of key information from contracts, aiding legal teams, contract managers, and businesses in streamlining contract analysis and compliance checks. This page provides you with a request body template for creating contract extractions, making it easy to integrate our service into your workflow.

## Why Use Extracta LABS for Contract Parsing?

* **Streamlined Contract Review**: Automate the extraction of key clauses, dates, and obligations.
* **Enhanced Compliance**: Easily identify and manage compliance requirements across contracts.
* **Customizable Data Extraction**: Focus on the specific details that matter most to your contract management process.

## Getting Started

To initiate contract parsing with Extracta LABS, you'll need to perform a `POST` request to our `/createExtraction` endpoint using a template tailored for contract documents. This template includes predefined keys that match typical contract elements, facilitating precise data extraction.

## Body Example for Contract

<details>

<summary>JSON Body</summary>

```json
{
    "extractionDetails": {
        "name": "Contract - Extraction",
        "language": "English",
        "fields": [
            {
                "key": "contract_details",
                "description": "contract details",
                "type": "object",
                "properties": [
                    {
                        "key": "contract_number",
                        "example": "56-2024",
                        "type": "string"
                    },
                    {
                        "key": "contract_date",
                        "example": "April 20, 2024",
                        "type": "string"
                    }
                ]
            },
            {
                "key": "involved_parties",
                "description": "information about the two parties involved in the contract",
                "type": "array",
                "items": {
                    "type": "object",
                    "properties": [
                        {
                            "key": "name",
                            "type": "string"
                        },
                        {
                            "key": "address",
                            "type": "string"
                        },
                        {
                            "key": "legal_representative",
                            "type": "string"
                        },
                        {
                            "key": "position",
                            "type": "string"
                        }
                    ]
                }
            },
            {
                "key": "contract_object",
                "description": "description of the contract object",
                "type": "string",
                "example": "Design and development of a web application as per the specifications in the attached Project Brief (Document A)"
            },
            {
                "key": "obligations",
                "description": "obligations of the involved parties",
                "type": "object",
                "properties": [
                    {
                        "key": "developer_obligations",
                        "example": "Design and develop a web application, deliver by September 1, 2024, make necessary adjustments within 30 days of delivery",
                        "type": "string"
                    },
                    {
                        "key": "client_obligations",
                        "example": "Pay the total fee in stages, test the application and provide feedback within 30 days of delivery",
                        "type": "string"
                    }
                ]
            },
            {
                "key": "contract_duration",
                "type": "string",
                "example": "12 months"
            },
            {
                "key": "contract_price",
                "type": "string",
                "example": "$200,000"
            }
        ]
    }
}
```

</details>

## How to Use This Template

1. **Generate an API Key**: If you haven’t already, sign up at [https://app.extracta.ai](https://app.extracta.ai/) and navigate to the `/api` page to create your API key.
2. **Execute the API Call**: Incorporate the JSON template in your `POST` request to `/createExtraction`. Remember to authenticate your request by including your API key as a Bearer token in the request header.
3. **Process the Extracted Data**: Upon submitting your request, the API will return structured data extracted from the contract, ready to be utilized in your systems or processes.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {}; // the json body from the example

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {} # the json body from the example

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}
{% endtabs %}

## Conclusion

With Extracta LABS, navigating through the complexities of contract analysis becomes a streamlined, automated process. Our API empowers you to focus on strategic aspects of contract management by handling the tedious task of data extraction.

For further details on integrating our API and making the most out of its capabilities, please refer to our [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction) and [1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction) pages.


# Business Card

Transform your contract management process with Extracta’s cutting-edge document parsing capabilities. Our API facilitates the extraction of key information from contracts, empowering legal teams, procurement departments, and businesses to automate data extraction, thereby enhancing contract analysis and management efficiency. This page offers a specialized request body template for contract parsing, designed to streamline your integration with our service.

## Why Extracta LABS for Contract Parsing?

* **Efficient Data Extraction**: Automate the extraction of critical contract elements, such as parties involved, terms, conditions, and obligations.
* **High Accuracy and Intelligence**: Utilize Extracta’s advanced algorithms for high-precision data extraction.
* **Customizable and Scalable**: Adapt the data extraction to suit specific business needs, scalable for organizations of any size.

## Getting Started

Initiate the process of parsing contracts through Extracta LABS by making a `POST` request to our `/createExtraction` endpoint. Below, we provide a request body template tailored for contract parsing. This template is preconfigured with keys for standard contract elements, streamlining the extraction process.

## Body Example for Business Card

<details>

<summary>JSON Body</summary>

```json
{
    "extractionDetails": {
        "name": "Business Card - Extraction",
        "language": "English",
        "fields": [
            {
                "key": "name",
                "description": "Name of the person in the business card",
                "example": "Bardahan Alexandru"
            },
            {
                "key": "job_title",
                "description": "Job title of the person",
                "example": "CEO"
            },
            {
                "key": "company_name",
                "description": "Extract company name from card; if absent, deduce from email, website, or social domains.",
                "example": "Fastapp Development"
            },
            {
                "key": "address",
                "description": "Address of the company",
                "example": "Bucharest, Romania"
            },
            {
                "key": "phone_numbers",
                "type": "array",
                "items": {
                    "type": "string",
                    "example": "+40 755 852 411"
                }
            },
            {
                "key": "email_addresses",
                "type": "array",
                "items": {
                    "type": "string",
                    "example": "office@fastappgroup.com"
                }
            },
            {
                "key": "website_url",
                "description": "Website url of the company",
                "example": "https://fastappgroup.com"
            },
            {
                "key": "social_media_handles",
                "description": "Social media handles of the company",
                "type": "array",
                "items": {
                    "type": "object",
                    "properties": [
                        {
                            "key": "type",
                            "description": "Type of the social media handle",
                            "example": "facebook"
                        },
                        {
                            "key": "username",
                            "description": "Username of the social media handle",
                            "example": "fastappgroup"
                        }
                    ]
                }
            }
        ]
    }
}
```

</details>

#### How to Use This Template

1. **Generate an API Key**: Sign up at [https://app.extracta.ai](https://app.extracta.ai/) and visit the `/api` page to create your API key.
2. **Execute the API Call**: Employ the JSON template in a `POST` request to `/createExtraction`, incorporating your API key in the request header as a Bearer token for authentication.
3. **Process the Extracted Data**: The API will respond with structured data extracted from the contract, ready for integration into your systems or workflows.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {}; // the json body from the example

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {} # the json body from the example

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}
{% endtabs %}

## Conclusion

Leverage Extracta LABS to elevate your contract management processes, from faster data retrieval to more informed decision-making. Our API not only simplifies the extraction of crucial contract details but also integrates seamlessly into various CRM platforms like Zoho, Salesforce, and HubSpot, enhancing your workflow automation.

For comprehensive instructions on integrating our API and leveraging its full capabilities, please refer to our [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction) and [1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction) pages.


# Email

Extracta LABS revolutionizes the way businesses handle email data through our Email Parsing API. By automating the extraction of text, attachments, and other relevant information from emails, we empower businesses to streamline their workflows, improve data management, and enhance business intelligence. This guide provides you with the essentials for integrating email parsing into your applications using our API.

## Why Extracta LABS for Email Parsing?

* **Comprehensive Data Extraction**: From body text to attachments, capture every piece of valuable information.
* **High Precision**: Benefit from Extracta’s advanced OCR and text processing technologies for accurate data extraction.
* **Custom Workflows**: Easily integrate extracted data into your CRM, database, or custom workflow for immediate use.

## Getting Started

Integrating email parsing into your system starts with a simple `POST` request to our `/createExtraction` endpoint, utilizing a request body designed for email data. Below, we provide a JSON template tailored for email parsing, highlighting the key information segments typically extracted from emails.

## Body Example for Email

<details>

<summary>JSON Body</summary>

```json
{
    "extractionDetails": {
        "name": "Email Data Extraction",
        "language": "English",
        "fields": [
            {
                "key": "email_info",
                "type": "object",
                "properties": [
                    {
                        "key": "Subject Line",
                        "description": "The subject of the email"
                    },
                    {
                        "key": "Email Date",
                        "description": "The date the email was sent",
                        "example": "February 20, 2024"
                    }
                ]
            },
            {
                "key": "sender_info",
                "type": "object",
                "properties": [
                    {
                        "key": "Sender Name",
                        "description": "The name of the person who sent the email"
                    },
                    {
                        "key": "Sender Email",
                        "description": "The email address of the sender"
                    },
                    {
                        "key": "Sender's Position",
                        "description": "The job title of the sender",
                        "example": "Project Manager"
                    },
                    {
                        "key": "Sender's Contact Number",
                        "description": "The phone number of the sender"
                    }
                ]
            },
            {
                "key": "recipient",
                "type": "object",
                "properties": [
                    {
                        "key": "Recipient Name",
                        "description": "The name of the email's recipient"
                    },
                    {
                        "key": "Recipient Email",
                        "description": "The email address of the recipient"
                    }
                ]
            },
            {
                "key": "cc",
                "description": "Carbon copy recipients of the email",
                "type": "array",
                "items": {
                    "type": "string",
                    "example": "test@gmail.com"
                }
            },
            {
                "key": "bcc",
                "description": "Blind carbon copy recipients of the email",
                "type": "array",
                "items": {
                    "type": "string",
                    "example": "test@gmail.com"
                }
            },
            {
                "key": "meeting_info",
                "description": "Information about the meeting",
                "type": "object",
                "properties": [
                    {
                        "key": "Meeting Date",
                        "example": "February 25, 2024"
                    },
                    {
                        "key": "Meeting Time",
                        "example": "10:00 AM"
                    },
                    {
                        "key": "Meeting Location",
                        "example": "Bucharest, Romania"
                    }
                ]
            },
            {
                "key": "project_name",
                "description": "The name of the project being discussed"
            },
            {
                "key": "proposal_equirements",
                "description": "Specific requirements or topics requested in the proposal",
                "type": "array",
                "items": {
                    "type": "string",
                    "example": "Integration of AI technologies"
                }
            }
        ]
    }
}
```

</details>

## How to Implement

1. **API Key Generation**: First, sign up at [https://app.extracta.ai](https://app.extracta.ai/) and create an API key on the `/api` page.
2. **API Request**: Use the JSON template in a `POST` request to `/createExtraction`, including your API key in the header for authentication.
3. **Data Integration**: The API's response will contain structured data extracted from the email, ready for integration into your system.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {}; // the json body from the example

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {} # the json body from the example

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}
{% endtabs %}

## Maximize Your Data Potential

With Extracta LABS, transforming unstructured email data into actionable insights has never been easier. Our Email Parsing API is a powerful tool for businesses looking to automate data extraction from emails for CRM updates, database enrichment, batch processing, or business intelligence analysis.

Explore further details on our API’s capabilities by visiting our [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction) and our [1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction) pages.


# Invoice

Transform the way you process invoices with Extracta's cutting-edge document parsing capabilities. Our API is specifically designed to automate the extraction of detailed invoice data, facilitating faster, more accurate accounts payable and receivable processes. This guide provides you with a tailored request body for creating invoice extractions, making it easier to integrate our technology into your financial workflows.

## Why Choose Extracta LABS for Invoice Parsing?

* **Efficient Processing**: Automate the extraction of invoice data to accelerate your financial operations.
* **High Precision**: Depend on our advanced OCR and machine learning technologies for accurate data capture.
* **Flexibility**: Extract data from invoices in various formats, including PDF, PNG, and JPG.
* **Custom Extraction**: Specify the data you need, from vendor details to line items, for tailored results.

## Getting Started

To initiate invoice data extraction, execute a `POST` request to the `/createExtraction` endpoint using the following request body template. This template is pre-configured with keys for common invoice data points, ensuring a streamlined setup for your extraction needs.

## Body Example for Invoice

<details>

<summary>JSON Body</summary>

````json
{
    "extractionDetails": {
        "name": "Invoice - Extraction",
        "language": "English",
        "options": {
            "hasTable": true,
            "handwrittenTextRecognition": false
        },
        "fields": [
            {
                "key": "invoice_id",
                "example": "1234567890",
                "type": "string"
            },
            {
                "key": "invoice_date",
                "description": "the invoice date in the following format yyyy-mm-dd",
                "example": "2022-01-01",
                "type": "string"
            },
            {
                "key": "invoice_tax_rate",
                "description": "the tax rate of the invoice",
                "example": "19%",
                "type": "string"
            },
            {
                "key": "invoice_currency",
                "description": "the currency of the invoice. return one of the following: RON, EUR, USD, LEI, etc.",
                "example": "Lei",
                "type": "string"
            },
            {
                "key": "client",
                "description": "the client in the invoice",
                "type": "object",
                "properties": [
                    {
                        "key": "client_name",
                        "description": "name of the client",
                        "example": "NovaTech Innovations",
                        "type": "string"
                    },
                    {
                        "key": "client_address",
                        "description": "address of the client",
                        "example": "1234 Main St, New York, NY 10001",
                        "type": "string"
                    },
                    {
                        "key": "client_tax_id",
                        "description": "tax id or vat id of the client",
                        "example": "987654321",
                        "type": "string"
                    }
                ]
            },
            {
                "key": "merchant",
                "description": "the merchant in the invoice",
                "type": "object",
                "properties": [
                    {
                        "key": "merchant_name",
                        "description": "name of the merchant",
                        "example": "Galactic Solutions",
                        "type": "string"
                    },
                    {
                        "key": "merchant_address",
                        "description": "address of the merchant",
                        "example": "789 Elm Rd, Seattle, WA 98109",
                        "type": "string"
                    },
                    {
                        "key": "merchant_tax_id",
                        "description": "tax id or vat id of the merchant",
                        "example": "123987456",
                        "type": "string"
                    }
                ]
            },
            {
                "key": "items",
                "description": "the items in the invoice",
                "type": "array",
                "items": {
                    "type": "object",
                    "properties": [
                        {
                            "key": "name",
                            "description": "the name of the item",
                            "example": "Item 1",
                            "type": "string"
                        },
                        {
                            "key": "quantity",
                            "description": "the quantity of the item",
                            "example": "1",
                            "type": "string"
                        },
                        {
                            "key": "unit_price",
                            "description": "the unit price of the item. return only the number as a string.",
                            "example": "100.00",
                            "type": "string"
                        },
                        {
                            "key": "total_price",
                            "description": "the total price of the item. return only the number as a string.",
                            "example": "100.00",
                            "type": "string"
                        },
                        {
                            "key": "tax",
                            "description": "the total tax of the item. return only the number as a string.",
                            "example": "100.00",
                            "type": "string"
                        }
                    ]
                }
            },
            {
                "key": "invoice_sub_total",
                "description": "The total amount before tax is applied. This may correspond to 'Total' or similar terms before tax addition. Return only the number as a string.",
                "example": "863.03",
                "type": "string"
            },
            {
                "key": "invoice_tax_amount",
                "description": "The amount of tax charged on the invoice. This is the tax figure listed separately from the sub-total. Return only the number as a string.",
                "example": "163.97",
                "type": "string"
            },
            {
                "key": "invoice_grand_total",
                "description": "The total amount after tax has been added. This should capture the final payable amount. Return only the number as a string.",
                "example": "1027.00",
                "type": "string"
            }
        ]
    },
    "file": "https://deveatery.com/extracta/invoice.png"
}
```
````

</details>

## How to Apply This Template

1. **Obtain an API Key**: Sign up at [https://app.extracta.ai](https://app.extracta.ai/) and visit the `/api` page to generate your unique API key.
2. **Craft Your API Request**: Incorporate the JSON template in a `POST` request to `/createExtraction`. Remember to authenticate your request by including your API key as a Bearer token in the request header.
3. **Process the Extracted Data**: Once the API processes your request, it will return structured data extracted from the invoice, ready for integration into your financial systems.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {}; // the json body from the example

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {} # the json body from the example

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}
{% endtabs %}

## Conclusion

Leverage Extracta LABS to modernize your invoice processing, reducing manual data entry and enhancing the accuracy of your financial operations. Our API simplifies the task of digitizing, capturing, and structuring data from scanned invoices, allowing you to focus on strategic financial management.

Explore our [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction) and [1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction) pages for more insights on integrating Extracta LABS into your document processing workflow.


# Receipt

Transform your receipt processing with Extracta’s cutting-edge document parsing capabilities. Our API is designed to automate the extraction of data from receipts, ideal for retailers, financial institutions, and any business that needs to analyze purchase data efficiently. This page guides you through using our API for receipt parsing, featuring a pre-configured request body template for immediate use.

## Why Extracta LABS for Receipt Parsing?

* **Streamline Financial Workflows**: Automate the capture and analysis of receipt data to speed up accounting processes and financial reporting.
* **Enhanced Accuracy**: Benefit from high precision OCR and data parsing technology that minimizes errors in financial analysis.
* **Custom Data Extraction**: Extract only the data you need, from transaction dates and amounts to item descriptions and merchant information.

## Getting Started

To extract data from receipts using Extracta LABS, you will need to send a `POST` request to our `/createExtraction` endpoint with a specific request body designed for receipt data. Below, we provide a template that includes predefined keys for common receipt information fields.

## Body Example for Receipt

<details>

<summary>JSON Body</summary>

```json
{
    "extractionDetails": {
        "name": "Receipt - Extraction",
        "language": "English",
        "options": {
            "hasTable": true,
            "handwrittenTextRecognition": false
        },
        "fields": [
            {
                "key": "receipt_id",
                "example": "1234567890",
                "type": "string"
            },
            {
                "key": "receipt_date",
                "description": "the invoice date in the following format yyyy-mm-dd",
                "example": "2022-01-01",
                "type": "string"
            },
            {
                "key": "merchant",
                "description": "the merchant in the invoice",
                "type": "object",
                "properties": [
                    {
                        "key": "merchant_name",
                        "description": "name of the merchant",
                        "example": "Galactic Solutions",
                        "type": "string"
                    },
                    {
                        "key": "merchant_address",
                        "description": "address of the merchant",
                        "example": "789 Elm Rd, Seattle, WA 98109",
                        "type": "string"
                    },
                    {
                        "key": "merchant_tax_id",
                        "description": "tax id or vat id of the merchant",
                        "example": "123987456",
                        "type": "string"
                    }
                ]
            },
            {
                "key": "items",
                "description": "the items in the receipt",
                "type": "array",
                "items": {
                    "type": "object",
                    "properties": [
                        {
                            "key": "name",
                            "example": "Item 1",
                            "type": "string"
                        },
                        {
                            "key": "quantity",
                            "example": "1",
                            "type": "string"
                        },
                        {
                            "key": "unit_price",
                            "description": "return only the number as a string.",
                            "example": "100.00",
                            "type": "string"
                        },
                        {
                            "key": "total_price",
                            "description": "return only the number as a string.",
                            "example": "100.00",
                            "type": "string"
                        }
                    ]
                }
            },
            {
                "key": "total_tax_amount",
                "description": "The amount of tax charged on the invoice. This is the tax figure listed separately from the sub-total. Return only the number as a string.",
                "example": "163.97",
                "type": "string"
            },
            {
                "key": "grand_total",
                "description": "The total amount after tax has been added. This should capture the final payable amount. Return only the number as a string.",
                "example": "1027.00",
                "type": "string"
            }
        ]
    }
}
```

</details>

## Utilizing the Template

1. **Obtain an API Key**: First, ensure you have an API key by registering at [https://app.extracta.ai](https://app.extracta.ai/) and creating your key on the `/api` page.
2. **Execute the API Request**: Incorporate the JSON template in your `POST` request to `/createExtraction`, including your API key in the header for authentication.
3. **Integrate Extracted Data**: The API will process your receipt and return detailed, structured data, ready to be integrated into your financial systems or applications.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {}; // the json body from the example

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {} # the json body from the example

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}
{% endtabs %}

## Leveraging Receipt Data

With Extracta LABS, you can unlock the potential of your receipt data, automating and refining financial processes with our advanced parsing technology. This not only saves time but also provides accurate, actionable insights from your transactional documents.

For more information on how to integrate Extracta LABS into your workflows, refer to our detailed  [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction) and [1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction) pages.


# Bank Statement

Transform the way you process financial documents with Extracta’s cutting-edge bank statement parsing capabilities. Our API is engineered to automate the extraction of financial data from bank statements, making it an indispensable tool for financial institutions, accounting software, and fintech startups. This guide provides a detailed overview of using a pre-defined request body for efficient bank statement parsing, ensuring a smooth integration with our platform.

## Why Choose Extracta LABS for Bank Statement Parsing?

* **Streamlined Data Extraction**: Automate the retrieval of transaction data, account balances, and other financial information.
* **OCR and AI-Powered**: Benefit from advanced OCR and AI technologies for high-precision reading of both digital and scanned bank statements.
* **Custom Extraction Paths**: Tailor the data extraction process to fit your specific requirements for financial analysis and reporting.

## Getting Started

To initiate the parsing of bank statements, you will need to execute a `POST` request to our `/createExtraction` endpoint. Below, we provide a request body template designed specifically for bank statement parsing. This template outlines predefined keys that are essential for extracting data from bank statements, such as transactions, account details, and summary.

## Body Example for Bank Statement

<details>

<summary>JSON Body</summary>

```json
{
    "extractionDetails": {
        "name": "Bank Check - Extraction",
        "language": "English",
        "fields": [
            {
                "key": "amount",
                "example": "25"
            },
            {
                "key": "amount_text",
                "example": "Twenty-five and no/100"
            },
            {
                "key": "bank_address"
            },
            {
                "key": "bank_name",
                "example": "WASHINGTON'S OLDEST NATIONAL BANK FIRST NATIONAL BANK"
            },
            {
                "key": "date",
                "example": "2019-10-13"
            },
            {
                "key": "memo",
                "example": "world hunger relief"
            },
            {
                "key": "payer_name",
                "example": "HON. GERALD R. FORD MRS. BETTY B. FORD"
            },
            {
                "key": "payer_address",
                "example": "1600 PENNSYLVANIA AVE. WASHINGTON, DC 20500"
            },
            {
                "key": "receiver_name",
                "example": "Presiding Bishop, Episcopal Church"
            },
            {
                "key": "receiver_address",
                "example": "815 Second Avenue, New York, NY 10017"
            },
            {
                "key": "currency",
                "example": "USD",
                "description": "currency as string, not symbol"
            },
            {
                "key": "micr_raw",
                "example": "A0540D0004A C140611 6C B0000002500B"
            },
            {
                "key": "account_number",
                "example": "6111611",
                "description": "Part of the MICR line, typically at the end."
            },
            {
                "key": "routing_number",
                "example": "05400004",
                "description": "Part of the MICR line, typically in the middle."
            },
            {
                "key": "check_number",
                "example": "878",
                "description": "Located at the top right of the check."
            }
        ]
    }
}
```

</details>

## How to Implement This Template

1. **Generate an API Key**: First, create an account on [https://app.extracta.ai](https://app.extracta.ai/) and visit the `/api` page to obtain your API key.
2. **API Call Execution**: Utilize the JSON template in your `POST` request to `/createExtraction`, incorporating your API key in the request header as a Bearer token.
3. **Data Integration**: Once the API processes your request, you will receive structured data from the bank statement, ready for integration into your system or workflow.

## Code Example

{% tabs %}
{% tab title="JavaScript" %}

```javascript
const axios = require('axios');

/**
 * Initiates a new document extraction process with the provided details.
 * 
 * @param {string} token - The authorization token for API access.
 * @param {Object} extractionDetails - The details of the extraction to be created.
 * @returns {Promise<Object>} The promise that resolves to the API response with the new extraction ID.
 */
async function createExtraction(token, extractionDetails) {
    const url = "https://api.extracta.ai/api/v1/createExtraction";

    try {
        const response = await axios.post(url, {
            extractionDetails
        }, {
            headers: {
                'Content-Type': 'application/json',
                'Authorization': `Bearer ${token}`
            }
        });

        // Handling response
        return response.data; // Directly return the parsed JSON response
    } catch (error) {
        // Handling errors
        throw error.response ? error.response.data : new Error('An unknown error occurred');
    }
}

async function main() {
    const token = 'apiKey';
    const extractionDetails = {}; // the json body from the example

    try {
        const response = await createExtraction(token, extractionDetails);
        console.log("New Extraction Created:", response);
    } catch (error) {
        console.error("Failed to create new extraction:", error);
    }
}

main();
```

{% endtab %}

{% tab title="Python" %}

```python
import requests


def create_extraction(token, extraction_details):
    url = "https://api.extracta.ai/api/v1/createExtraction"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {token}"}

    try:
        response = requests.post(url, json=extraction_details, headers=headers)
        response.raise_for_status()  # Raises an HTTPError if the response status code is 4XX/5XX
        return response.json()  # Returns the parsed JSON response
    except requests.RequestException as e:
        # Handles any requests-related errors
        print(e)
        return None


# Example usage
if __name__ == "__main__":
    token = "apiKey"
    extraction_details = {} # the json body from the example

    response = create_extraction(token, extraction_details)
    print("New Extraction Created:", response)
```

{% endtab %}
{% endtabs %}

## Conclusion

With Extracta LABS, parsing bank statements becomes an automated, accurate, and efficient part of your financial document processing workflow. Our API leverages the latest in OCR and AI technologies to provide you with the data you need when you need it, ensuring your financial analyses are always informed by the most up-to-date information.

Explore further integration details and maximize the potential of our API by visiting our  [API Endpoints - Data Extraction](/data-extraction-api/api-endpoints-data-extraction) and [1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction) pages.


# Tutorials

Welcome to the Extracta.ai tutorials page! Here, we provide step-by-step guides to help you get started with our API. Whether you're new to Extracta.ai or looking to explore more advanced features, these tutorials are designed to equip you with the knowledge to integrate our API seamlessly into your workflows.

## Getting Started

### How to Make an Account and Obtain an API Key

1. **Sign Up**: Visit [https://app.extracta.ai](https://app.extracta.ai/) and sign up for an account to get started.
2. **Generate API Key**: Once logged in, navigate to the `/api` section of your dashboard. Here, you can generate a new API key by clicking on the "Generate API Key" button.
3. **Keep Your API Key Secure**: Store your API key in a safe place. It's your unique identifier and grants access to the API's capabilities.

{% content-ref url="/pages/nVOqzbgiuFMfYJFfSxNq" %}
[Authentication](/api-reference/authentication)
{% endcontent-ref %}

### First API Call: Creating an Extraction

Learn how to make your first API call to create an extraction template. This tutorial covers the basics of forming a request to our `/createExtraction` endpoint.

{% content-ref url="/pages/Kr8TtNYErc3zJQQ68V7K" %}
[1. Create extraction](/data-extraction-api/api-endpoints-data-extraction/1.-create-extraction)
{% endcontent-ref %}

## Postman Collection

For a complete and interactive set of API requests, please refer to our [Postman Integration](/data-extraction-api/postman-integration) collection.

## Document Types

### Parsing a Resume

This tutorial walks you through the process of extracting data from a resume using a predefined template. It includes steps for preparing the request and handling the extracted data.

{% content-ref url="/pages/QpnjK5EoHo2OnV9R5HLN" %}
[Resume / CV](/documents/resume-cv)
{% endcontent-ref %}

### Working with Custom Documents

Discover how to parse custom documents by specifying your own fields for extraction. This guide provides insights into customizing your extraction request for various document types.

{% content-ref url="/pages/zInUxfskFjgVjNXxdAy5" %}
[Custom Document](/documents/custom-document)
{% endcontent-ref %}

## Advanced Topics

### Batch Processing

Learn how to upload and process documents in batches, a feature that enhances efficiency by handling multiple documents in a single request.

{% content-ref url="/pages/sWhIeG7Es4VBfiHRfeYw" %}
[5. Upload Files](/data-extraction-api/api-endpoints-data-extraction/5.-upload-files)
{% endcontent-ref %}

### Receiving Batch Results: Polling vs. Webhook

This tutorial explains the two methods of receiving batch results: polling and using a webhook for real-time updates. Understand the pros and cons of each method and how to implement them.

{% content-ref url="/pages/xXq6iSQKgPG6jmHJXdgA" %}
[Polling vs Webhook](/data-extraction-api/polling-vs-webhook)
{% endcontent-ref %}

## FAQ

Find answers to common questions about using the Extracta.ai API, troubleshooting, and best practices.

{% content-ref url="/pages/dhCn1MEQki1HiLNMzR52" %}
[FAQ](/contact/faq)
{% endcontent-ref %}

## Need More Help?

If you encounter any issues or have further questions, our support team is here to assist you. Visit our Contact page for more information.

{% content-ref url="/pages/dpJM2AKMBKrLy4SW1M9M" %}
[Contact Us](/contact/contact-us)
{% endcontent-ref %}


# Contact Us

At Extracta LABS, we are committed to providing exceptional support and ensuring that your experience with our product is as seamless and productive as possible. Whether you have questions, feedback, or need assistance, our team is here to help.

## Get in Touch

### Email Support

For support or any inquiries, please email us at:

* **Email:** <office@extracta.ai>

### Phone Support

For immediate assistance, you can reach us by phone during business hours:

* **Phone:** +40 755 852 411
* **Hours:** Monday to Friday, 9 AM to 5 PM (GMT)

### Live Chat

Visit our website and use the live chat feature for real-time support from our team. Available during business hours.

<https://extracta.ai/>

### Book a demo

Visit our website to book a 30-minute personalized demo and explore how Extracta LABS can revolutionize your data extraction processes.

<https://extracta.ai/book-a-demo>

### Social Media

Follow us on our social media channels to stay updated on the latest news, updates, and announcements:

* **LinkedIn:** <https://www.linkedin.com/company/extracta-ai>
* **Reddit:** <https://www.reddit.com/r/Extracta/>
* **Discord:** <https://discord.gg/GQCjkvnqE6>


# FAQ

Welcome to the Extracta LABS FAQ page! Here, you'll find answers to some of the most common questions about our Document Parsing API, how to use it, and what you can expect. If your question isn't answered here, please don't hesitate to contact us.

<details>

<summary>What is Extracta LABS?</summary>

Extracta LABS is a cutting-edge technology platform that specialises in extracting structured data from any type of document (like CVs and invoices). Our service is designed to automate workflows, replace manual data processing, and enhance efficiency across various industries.

</details>

<details>

<summary>What types of documents can we process?</summary>

We can process a wide range of documents both structured and unstructured, including PDFs, Word documents, text files, and scanned images (PNG, JPG) using OCR technology where necessary.

</details>

<details>

<summary>Is our platform GDPR compliant?</summary>

Absolutely. Both our web platform and API are fully GDPR compliant. We prioritize data privacy and security in all our operations.

</details>

<details>

<summary>Can it be integrated into existing systems?</summary>

Yes, Extracta LABS is designed for seamless integration. You can connect our service to your existing software and workflows through our API. Additionally, in the future we plan to offer deployment on local systems for enhanced data privacy.

</details>

<details>

<summary>How do we differ from our competitors?</summary>

Unlike competitors that rely on predefined templates and models, Extracta LABS uses fine-tuned LLMs to extract data from any document without prior training, offering up to 99% accuracy. This allows for greater flexibility and faster implementation at lower costs.

</details>

<details>

<summary>What pricing models do we offer?</summary>

We offer our solution on a Pay-per-request basis, ensuring flexibility and scalability for our clients. We also provide a free trial to test our platform before committing.

</details>

<details>

<summary>Where can I find technical support or have other questions answered?</summary>

Our dedicated support team is available to assist you with any technical queries or further information. Please reach out to us through our website's contact form or support chat. Also, feel free to check our tutorials and documentation.

</details>


