Skip to content

Bulk user import ​

AuthHero implements Auth0's bulk user-import job API, so a migration script written against Auth0 works against AuthHero by changing only the base URL:

  • POST /api/v2/jobs/users-imports — submit a file of users
  • GET /api/v2/jobs/{id} — poll the job
  • GET /api/v2/jobs/{id}/errors — read per-user failures

The auth0 npm client's jobs.importUsers(), jobs.get() and jobs.errors() work unmodified.

Use this when you already hold your users' password hashes and want to own the records upfront. If you cannot extract hashes, use lazy migration instead — it drains the upstream tenant incrementally and needs no hashes at all. The two are complementary: importing profiles in bulk and leaving import_mode on for passwords is a perfectly good combination.

Submitting a job ​

bash
curl -X POST https://auth.example.com/api/v2/jobs/users-imports \
  -H "Authorization: Bearer $TOKEN" \
  -F [email protected] \
  -F connection_id=con_abc123 \
  -F upsert=true \
  -F external_id=batch-0042
FieldNotes
usersJSON file: an array of user objects (see below). Max 500 KB by default.
connection_idThe database connection to import into. Required.
upsertWhen true, update users matching on user_id, email, username or phone_number. Defaults to false, which makes an existing user a per-row error.
external_idYour own correlation id, echoed back on the job.
send_completion_emailAccepted for compatibility; not yet delivered.

The response is 202 with the job:

json
{
  "id": "job_op_3kf9...",
  "type": "users_import",
  "status": "pending",
  "connection_id": "con_abc123",
  "external_id": "batch-0042",
  "percentage_done": 0,
  "summary": { "total": 1000, "inserted": 0, "updated": 0, "failed": 0 }
}

Poll GET /api/v2/jobs/{id} until status is completed or failed. percentage_done and time_left_seconds are estimated from actual throughput.

The import file ​

json
[
  {
    "email": "[email protected]",
    "email_verified": true,
    "user_id": "abc123",
    "username": "jane",
    "name": "Jane Doe",
    "app_metadata": { "plan": "pro" },
    "user_metadata": { "theme": "dark" },
    "custom_password_hash": {
      "algorithm": "bcrypt",
      "hash": {
        "value": "$2b$10$C9hB01.YxRSTcn/ZOOo4j.TW7xCKKFKBSF.C7E0xiUwumqIDqWUXG"
      }
    }
  }
]

email is required. Profile fields (username, name, given_name, family_name, nickname, picture, phone_number, phone_verified, blocked), app_metadata and user_metadata are supported. Unknown fields are ignored rather than rejected.

A bare user_id is prefixed with the tenant's database provider, exactly as Auth0 prefixes with auth0|. Supplying abc123 stores auth0|abc123; supplying auth0|abc123 is not double-prefixed. Omit it and one is generated.

Passwords: bcrypt only ​

WARNING

AuthHero verifies passwords with bcrypt, so bcrypt is the only hash algorithm it can import.

Both Auth0 forms are accepted:

  • password_hash — a bcrypt string, $2a$ or $2b$
  • custom_password_hash — { "algorithm": "bcrypt", "hash": { "value": "$2b$..." } }, where $2a$, $2b$ and $2y$ are all accepted and hash.encoding must be utf8 if given

Any other algorithm (argon2, pbkdf2, scrypt, sha*, md5, hmac, ldap) fails that row only, with code UNSUPPORTED_HASH_ALGORITHM. The rest of the file still imports. This is deliberate: storing a hash AuthHero cannot verify would create a user who can never log in, which is worse than a clear error.

Users whose hashes cannot be imported — and users with no password field at all — are still created as valid accounts. They sign in through a password reset, or through lazy migration's upstream password fallback if you have it enabled. Once a user has a local password, that fallback is skipped, matching Auth0's behaviour of never re-delegating after import.

Reading errors ​

bash
curl https://auth.example.com/api/v2/jobs/job_op_3kf9.../errors \
  -H "Authorization: Bearer $TOKEN"

Returns 200 with an array, or 204 when the job produced no errors:

json
[
  {
    "user": { "email": "[email protected]", "password_hash": "[redacted]" },
    "errors": [
      {
        "code": "UNSUPPORTED_HASH_ALGORITHM",
        "message": "AuthHero can only import bcrypt password hashes; received \"argon2\".",
        "path": "custom_password_hash.algorithm"
      }
    ]
  }
]

The submitted user object is echoed back with credential material redacted — password hashes are never returned over the API.

CodeMeaning
VALIDATION_ERRORThe entry did not match the schema.
UNSUPPORTED_HASH_ALGORITHMA valid hash in an algorithm AuthHero cannot verify.
UNSUPPORTED_HASH_FORMATbcrypt, but a variant ($2$, $2x$) or encoding that would never verify.
DUPLICATE_ENTRYThe same identity appears earlier in the same file.
USER_ALREADY_EXISTSThe user exists and upsert was not enabled.
INTERNAL_ERRORThe write itself failed.

Importing millions of users ​

Jobs are durable and resumable, which is what makes large migrations safe.

Every user in the file is staged as its own database row before the request returns. Rows are then processed in chunks, and each chunk commits its outcomes before the next begins. If the process handling a job dies — evicted, redeployed, timed out — nothing is lost and nothing is done twice: the unprocessed rows are still marked pending, and the next driver resumes from the last committed chunk.

To guarantee that resumption happens, wire the sweep into your scheduled handler:

ts
import { resumeUsersImports, runRetention } from "authhero";

export default {
  async scheduled(_event, env) {
    // Picks up any import whose driver died, and carries it to completion.
    await resumeUsersImports(dataAdapter);

    // Deletes finished jobs and their staged rows after 24 hours.
    await runRetention({ dataAdapter });
  },
};

Without this the accepting request still makes a start, and any subsequent request or sweep will finish the job — but the scheduled sweep is what bounds how long a stalled job can sit.

Limits ​

LimitDefaultConfigurable
File size500 KB (~1,000 users)init({ usersImportMaxBytes })
Concurrent jobs per tenant2init({ usersImportMaxConcurrentJobs })
Job data retention24 hoursrunRetention({ usersImportRetentionHours })

The defaults are Auth0's, so an Auth0-shaped client sees identical behaviour. A migration of a million users is therefore ~1,000 jobs of ~1,000 users, submitted two at a time — the same shape as the equivalent Auth0 migration. Raise the limits only for a migration you control end to end; a larger file means more rows staged inside a single request.

Exceeding the concurrency limit returns 429. AuthHero does not queue past the limit, and neither does Auth0 — the client is expected to hold back its own submissions, which is why Auth0's own guidance for ten or more jobs is to drive them from a job-scheduler framework. A submitted-but-unfinished job counts toward the limit, not just an actively processing one.

How rows are written ​

Within a chunk, rows are batched rather than processed one at a time. The chunk is parsed and mapped in memory, each existence probe runs as one query per field across the whole chunk, and rows that turn out to be new users are written with a single batched insert. A 50-row chunk costs about five queries rather than about two hundred.

This matters more than it might sound. The work is dominated by round-trip latency, not by the database's write cost — against a hosted database a row-at-a-time loop spends roughly 400 ms per row almost entirely waiting, which is the difference between a million-user import taking hours and taking days.

Three cases deliberately keep the per-row path:

  • Upserts. A job submitted with upsert=true never batches at all: it resolves each row's identity immediately before writing it. An upsert changes rows that a later row in the same chunk may itself match — renaming a username frees that username for the next row — so a chunk-wide identity snapshot would not be equivalent to writing the rows in order. A bulk migration is overwhelmingly upsert: false.
  • A failed batch. Because a batch insert cannot say which row collided, any failure is retried row by row so each outcome is still attributed to the row that caused it.
  • Resuming an interrupted job. Rows whose derived id already exists are recognised as their own earlier write rather than reported as a conflict.

De-duplication within a chunk is handled explicitly, on every identifier the probes cover: two rows carrying the same email — or the same username, or the same phone number — produce one user, not two.

Adapters may implement an optional createMany on UserDataAdapter to take part in this; the kysely adapter does. It is optional, and AuthHero falls back to looping the per-row write when it is absent, so an adapter without it keeps working — just at the older cost per row. createMany writes the user and its password only, so users carrying identities, activity counters or outbox events continue to go through create.

Imported users do not run registration hooks ​

A fresh user is written with users.rawCreate, so pre-registration denial, hook metadata mutation, built-in email linking and the post-registration outbox event do not fire for imported users. This matches Auth0, where a users-import job does not trigger Actions — an import is a data-loading path, not a sign-up.

Both the batched write and the per-row fallback use rawCreate, deliberately. createMany is not wrapped by the hook layer, so routing one path through the hooks and the other around them would make the policy applied to an import file depend on which adapter happened to be installed.

If you need hooks to run for imported users, do it after the job completes — read the operation's rows and drive them through your own flow — rather than relying on the import to fire them.

Passwords commit with their user ​

An imported password is written in the same call as the user, and commits in the same transaction. An interrupted driver therefore cannot leave an imported user who exists but cannot log in.

Required scopes ​

EndpointScope
POST /api/v2/jobs/users-importscreate:users
GET /api/v2/jobs/{id}create:users or read:users
GET /api/v2/jobs/{id}/errorscreate:users or read:users

Setting it up ​

There is no workflow engine, queue, or object storage to configure. Jobs are durable because their progress is checkpointed in the database, so the only moving parts are your database schema and a cron.

A service already using AuthHero needs three changes.

1. Run the migrations (required) ​

The feature adds a tenant_operation_rows table and four columns on tenant_operations (input, result, claimed_by, claim_expires_at).

  • kysely — included in migrateToLatest(); run your usual migration step.
  • drizzle — the tables are control-plane, so apply the drizzle-control-plane/ migration set, not the core drizzle/ one.

2. Make sure the adapters are present (required) ​

The endpoints need the tenantOperations and tenantOperationRows adapters. Without them they return 501.

  • kysely — registered automatically, nothing to do.

  • drizzle — only registered when you pass controlPlane: true:

    ts
    const data = createAdapters(db, {
      useTransactions: true,
      controlPlane: true,
    });

Single-database and control-plane deployments only

Job records live in the control-plane tables, so bulk import is available where those tables and the users being imported share a database — a single-database deployment (for example kysely on PlanetScale), or the control plane itself.

It is not available on Workers-for-Platforms tenant workers, whose D1s carry only the core schema; those return 501. Import into such a tenant from the control plane, or use lazy migration.

3. Wire the sweeps into your scheduled handler (required) ​

ts
import { resumeUsersImports, runRetention } from "authhero";

export default {
  async scheduled(_event, env) {
    // Picks up any import whose driver died and carries it to completion.
    // This is what makes a large import reliable — without it, a job whose
    // driver was evicted mid-run waits for the next request to nudge it.
    await resumeUsersImports(dataAdapter);

    // Also deletes finished jobs and their staged rows after 24 hours.
    await runRetention({ dataAdapter });
  },
};

Every 1–5 minutes is a reasonable cadence. resumeUsersImports() returns { scanned, advanced, completed, errors } if you want to log it, takes the same budget options as a single pass (maxRows, chunkSize, deadline), and is safe to run concurrently with itself — a job already being advanced is left to its current driver.

If you already call runRetention(), the 24-hour job cleanup is picked up automatically with no code change.

4. Optionally raise the limits ​

Only if you are running a migration you control end to end:

ts
const { app } = init({
  dataAdapter,
  usersImportMaxBytes: 5 * 1024 * 1024, // default 500 KB
  usersImportMaxConcurrentJobs: 8, // default 2
});

Both default to Auth0's values, so leaving them alone keeps behaviour identical to Auth0 for any client.

Dual-licensed: AGPL-3.0-only or commercial license.