Database Integration Throughput
When Transcend processes a data subject request (DSR) against a database integration, it runs one SQL job per datapoint through your Sombra gateway. Two settings decide how fast those jobs finish: whether Sombra reuses database connections (connection pooling), and the Transcend rate limits on the integration. This guide explains how each one affects throughput and how to tune them.
- Transcend schedules a job for each datapoint a request needs to read or erase.
- The integration's rate limits decide how many of those jobs Transcend sends to Sombra each minute and each hour.
- Sombra gets a database connection, runs the SQL statement, and returns the result to Transcend.
Without pooling, Sombra opens a new connection for every job and closes it afterwards. Opening a connection involves a network handshake, TLS negotiation, and authentication, which often takes longer than the query itself. With pooling, Sombra keeps connections open and hands an existing one to the next job.
Pooling is a Sombra setting, so whether you can turn it on depends on how your Sombra is hosted:
- Self-hosted Sombra: set the environment variables below on your Sombra deployment.
- Dedicated Sombra hosted by Transcend: contact Transcend Support to enable pooling.
- Shared multi-tenant Sombra: connections are not pooled.
These settings apply to database integrations that connect over ODBC, such as PostgreSQL, MySQL, MariaDB, Snowflake, Amazon Redshift, Oracle, and IBM DB2. Microsoft SQL Server and Azure Synapse use a separate pool configured with the MSSQL_POOL_* environment variables, listed in the Environment Variables Reference.
Turn pooling on by setting ODBC_POOL_CACHE_SIZE on your Sombra, then redeploy Sombra. Connection Pooling lists every pooling setting, its default, and a starting point for high-volume workloads.
Each Sombra instance (container) keeps its own pools. For each connection string, your database can see up to ODBC_CONNECTION_POOL_SIZE connections from every Sombra instance. Instead of a constant stream of short-lived connections, the database sees a steady set of long-lived ones.
Before raising pool sizes, check that your database's maximum connection limit, and any per-user limit on the database user Sombra connects as, leave room for every Sombra instance. If you set ODBC_CONNECTION_POOL_SIZE_SHRINK to false, connections opened during busy periods stay open until the pool expires.
In the Admin Dashboard, open the database integration and go to the Connection tab. When the integration routes through a self-hosted or dedicated Sombra that does not have pooling turned on, a banner recommends enabling it. If the Sombra is too old to report its pooling settings, the banner asks you to upgrade Sombra instead.
You can also search your Sombra logs for these messages:
| Log message | What it means |
|---|---|
Creating new pool for connection endpoint | Sombra created a pool for a database. Pooling is on. |
Returning pooled <driver> connection for endpoint | A job reused a pooled connection. Pooling is working. |
Creating new <driver> connection for endpoint | Logged for every job, with no pool messages: pooling is off. |
Pooling makes each job faster; rate limits cap how many jobs Transcend sends. Database integrations start with Transcend-defined limits of 100 requests per minute and 5,000 requests per hour. These limits are a conservative ceiling that protects Sombra and your database when many integrations run at once, not a throughput target.
Turn on pooling first, then raise the limits if requests are still waiting on them. Raising limits without pooling sends more short-lived connections to Sombra and your database, which can slow processing down. To check whether limits are holding requests back and how to change them, see Data System Rate Limits.
One customer processing a high volume of DSRs against a single database saw these changes after enabling pooling with the high-volume starting point from Connection Pooling:
| Metric | Before pooling | After pooling |
|---|---|---|
| Jobs per hour | 160,000 to 190,000 | 475,000 to 560,000 |
| Average time per job | About 600 ms | 250 to 330 ms |
| Median database round trip | About 190 ms | About 26 ms |
| 95th percentile database round trip | About 1.2 s | About 61 ms |
Your results depend on query cost, network latency between Sombra and the database, and the database's own capacity.