# Load pandas data frame to neo4j database in batches

**URL:** <https://community.neo4j.com/t/load-pandas-data-frame-to-neo4j-database-in-batches/64221>\
**Category:** Import / Export\
**Created:** [September 21, 2023, 5:02pm UTC](https://community.neo4j.com/t/load-pandas-data-frame-to-neo4j-database-in-batches/64221 "2023-09-21T17:02:48Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![kamalika.ray](https://avatars.discourse-cdn.com/v4/letter/k/57b2e6/32.png) [@kamalika.ray](https://community.neo4j.com/u/kamalika.ray)\
**Post date:** [September 21, 2023, 5:02pm UTC](https://community.neo4j.com/t/load-pandas-data-frame-to-neo4j-database-in-batches/64221/1 "2023-09-21T17:02:48Z")

</div>

Hi,

I have a big sized pandas data frame. I want to load it into neo4j database using the neo4j python driver. How can I load it in batches?

OR

Is there any way I can use the file (stored in a local driver) directly something like this using the neo4j python driver:

LOAD CSV WITH HEADERS FROM 'file://///genes.csv' AS line CALL { WITH line CREATE (:Gene {symbol: line.symbol})} IN TRANSACTIONS OF 100 ROWS"

Thanks

---

<div class="post-metadata">

**Author:** ![rouven\_bauer](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/rouven_bauer/32/16547_2.png) [@rouven\_bauer](https://community.neo4j.com/u/rouven_bauer)\
**Post date:** [September 25, 2023, 2:03pm UTC](https://community.neo4j.com/t/load-pandas-data-frame-to-neo4j-database-in-batches/64221/2 "2023-09-25T14:03:42Z")

</div>

Hi,

if you want to import CSV files, they must either be located in the import directory on the server or be accessible through HTTP(S), or FTP. See also [the docs](https://neo4j.com/docs/getting-started/data-import/csv-import/).

As to how to batch things manually, here's a suggestion I threw together quickly. Please play around with it, adjust it to your needs, and fine-tune the constants. Especially the batch-size is very dependent on the work load.

```python
import asyncio

import neo4j
import numpy as np
import pandas as pd

# some sample data
data = np.stack(
    np.meshgrid(np.arange(-1000, 1000), np.arange(-1000, 1000)), -1
).reshape(-1, 2)
df = pd.DataFrame(data, columns=["x", "y"])

URL = "neo4j://localhost:7687"
AUTH = ("neo4j", "pass")
DB = "neo4j"
BATCH_SIZE = 10000
MAX_CONCURRENCY = 50

async def upload_batch(semaphore, driver, batch):
    async with semaphore:
        await driver.execute_query(
            "UNWIND range(0, $len - 1) AS i "
            "CREATE (n:Node {x: $data['x'][i], y: $data['y'][i]})",
            data=batch,
            len=len(batch),
            database_=DB,
        )

async def main():
    semaphore = asyncio.Semaphore(MAX_CONCURRENCY)

    async with neo4j.AsyncGraphDatabase.driver(URL, auth=AUTH) as driver:
        # optionally clear out the DB for testing
        # await driver.execute_query("MATCH (n) DETACH DELETE n", database_=DB)
        tasks = [
            upload_batch(semaphore, driver, df[offset:(offset + BATCH_SIZE)])
            for offset in range(0, len(df), BATCH_SIZE)
        ]
        await asyncio.gather(*tasks)

if __name__ == " __main__":
    asyncio.run(main())

```
