COPY INTO vs Auto Loader vs UI Upload: The Five Ways Data Gets Into Databricks
Databricks Data Analyst · Trade-off

COPY INTO vs Auto Loader vs UI Upload: The Five Ways Data Gets Into Databricks

Data gets into Databricks by one of five paths: the UI upload for a one-off file, COPY INTO for a bounded set of files you may reload, Auto Loader as a streaming table for files that keep arriving, Lakeflow Connect for SaaS applications and databases, and Delta Sharing or Marketplace for data another organization maintains. Two questions pick the path: who owns the data, and does it arrive once or keep arriving.

Last updated September 2026.

The short version

If the data is yours and arrives once, upload it or COPY INTO it. If it is yours and keeps arriving, Auto Loader behind a streaming table. If it lives in an application or database, a Lakeflow Connect managed connector. If someone else maintains it, share it with Delta Sharing or install it from Marketplace, and never copy what another organization keeps current.

Your own files: upload, COPY INTO, Auto Loader

UI upload

The Create or modify table dialog takes a file from your machine, infers a schema you can override in the preview (turn that zip code back into a string before it loses its leading zeros), and creates a managed Delta table in the catalog and schema you pick. It can also append to an existing table. It is the right path exactly once per file: a spreadsheet from a manager, a reference list, a sample to explore.

COPY INTO

COPY INTO finance.bronze.prices FROM 's3://vendor/prices/' FILEFORMAT = PARQUET loads the files at a path into an existing Delta table and, crucially, is idempotent: files it has already loaded are skipped, so re-running after a failure is safe, and a specific file can be forced to reload. It is the fit for a bounded set, a few thousand files at most, that arrives on a cadence you control: the quarterly drop, the monthly extract.

Auto Loader as a streaming table

CREATE STREAMING TABLE bronze.events AS SELECT * FROM STREAM read_files('s3://bucket/events/') is Auto Loader from Databricks SQL. It discovers new files as they land, processes each one exactly once, evolves the schema when new columns appear, and scales to millions of files without listing the folder every time. It is the path for data that keeps arriving: clickstream, sensor drops, an application that writes a file every few minutes. The same feature runs inside Lakeflow Spark Declarative Pipelines for engineering teams.

Someone else's data: Lakeflow Connect, Delta Sharing, Marketplace

Lakeflow Connect

Managed connectors for SaaS applications and databases: point-and-click pipelines that pull from the source's API or change feed on a schedule and write streaming tables, with the credentials held in a Unity Catalog connection. Partner Connect provisions Fivetran and other ingestion partners the same way. Writing your own API code against the Files API and SDKs is the last resort when no connector exists.

Delta Sharing

A provider creates a share, adds tables, and grants a recipient. Databricks-to-Databricks recipients mount the share as a read-only catalog with no tokens to manage; open-protocol recipients get a credential file and read through Spark, pandas or a BI connector. Nothing is copied: the provider's changes appear in place, and access ends when the share is revoked. The 2026 documentation calls the feature OpenSharing; the exam guide says Delta Sharing.

Marketplace

Delta Sharing with a storefront. Find the listing, get access, name the catalog, query. The provider keeps the data current, which is why the monthly CSV download from a vendor's website is the wrong answer whenever a Marketplace listing exists.

Side by side

Signal in the scenarioPathWhy
"A manager sends a spreadsheet"UI uploadOne file, once, schema fixed in the preview
"A few thousand files, reload occasionally"COPY INTOIdempotent, bounded, easy to force a reload
"New files every few minutes, millions over time"Auto Loader (streaming table)Incremental, exactly once, schema evolution
"It lives in Salesforce" or "in our Postgres"Lakeflow ConnectA managed connector beats an API you call yourself
"A partner keeps their tables current for us"Delta SharingNever copy what someone else maintains
"We want the census data"MarketplaceA listing, not a pipeline

Sourced from the Databricks documentation on ingestion, COPY INTO, Auto Loader, streaming tables in Databricks SQL, Lakeflow Connect and OpenSharing.

What this means for the exam

Exam tip: Databricks' own guidance is to start with the most managed layer that fits: a connector before a pipeline, a pipeline before hand-written code. When two options both work, the exam wants the one that copies less and needs less babysitting.

The common distractors are COPY INTO offered for a folder that never stops growing (it works, but Auto Loader is cheaper and incremental), a streaming table offered for a one-off spreadsheet, and any answer that downloads or re-exports data a share would have delivered in place. Federation is not on this list on purpose: it queries a remote database without ingesting it, which is a different question.

FAQ

What is the difference between COPY INTO and Auto Loader in Databricks?

COPY INTO loads a bounded set of files from a path into a Delta table and skips files it has already loaded, so it is idempotent and easy to re-run. Auto Loader, used as a streaming table, discovers and processes new files incrementally as they arrive and scales to millions of files with schema evolution.

When should I use the Databricks UI to upload a file?

For a one-off file such as a spreadsheet from a colleague. The upload dialog infers the schema, lets you override column types before the table is created, and can append to an existing table, but it is a manual path, not a pipeline.

How do I get data from Salesforce or a database into Databricks?

Use a Lakeflow Connect managed connector, which pulls from the source on a schedule into streaming tables with credentials held in a Unity Catalog connection. Partner Connect provisions third-party ingestion tools the same way, and custom API code is the last resort.

Should I copy a partner’s data into Databricks or share it?

Share it. Delta Sharing (called OpenSharing in the 2026 docs) mounts the provider’s tables as a read-only catalog with no copy, so their updates appear in place and access ends when the share is revoked. Marketplace is the same mechanism for published data products.

Practice this hands-on

The five ingestion paths are a trade-off chapter in the Data Analyst Associate course, followed by a chapter on the UI upload itself.

Open the chapter →