How to Evaluate an ETL Platform: What Separates the Tools

ETL and ELT platforms separate mainly on connector breadth, whether transformation runs in the pipeline tool itself or gets pushed down into the warehouse, and whether the tool covers extract-and-load, transformation, or both.

What does an ETL platform do in a modern data stack?

Most vendors in this category today run ELT rather than classic ETL: they extract data from source systems, load it into a warehouse largely as-is, and leave transformation to run inside the warehouse using its own compute. That shift matters for evaluation because it splits the category into tools built for the extract-and-load half, tools built for the transformation half, and a smaller set that combine pieces of both.

Buyers typically start this evaluation because a marketing or data team is maintaining custom scripts to move data between ad platforms, CRMs, and a warehouse, and that maintenance burden has become the bottleneck rather than the analysis work itself.

What separates the tools when you get into evaluation?

Connector breadth. Airbyte ships more than 350 pre-built connectors plus a Connector Builder and CDK for building custom ones, and handles schema evolution on source changes automatically. Matillion covers roughly 100 connectors through its Data Loader and Hub catalog, a narrower set by comparison, with its investment concentrated instead on the transformation and orchestration side. dbt Labs does not extract or load data at all; connector breadth is out of scope for the product by design, since dbt operates purely as a transformation layer on data that is already in the warehouse.

Where transformation runs. Matillion pushes transformation logic down into the destination warehouse's own compute, supporting Snowflake, BigQuery, Redshift, and Delta Lake as destinations, with a low-code visual interface for building those pipelines. Airbyte handles extract-and-load and integrates natively with dbt for the transformation step rather than building deep transformation tooling of its own, effectively pairing with a tool like dbt rather than competing with it. dbt itself is the transformation layer many other tools integrate with: it handles model materializations, incremental build strategies, and macros, with adapters covering Snowflake, BigQuery, Databricks, Redshift, and DuckDB as execution targets.

Orchestration model. Airbyte delegates broader workflow orchestration to partner tools like Airflow, Dagster, or Prefect rather than building an orchestration layer natively, which means teams with complex multi-step pipelines need to plan for that additional tool. Matillion builds orchestration into its own pipelines directly. dbt Cloud includes a managed scheduler with dependency DAGs for transformation jobs, and dbt Mesh extends that into cross-project lineage and access control for organizations running multiple dbt projects.

Deployment and pricing model. Airbyte offers both a usage- and credits-based cloud product and a self-hosted open-source version; self-hosting removes licensing cost but adds the overhead of running and tuning the infrastructure, typically Kubernetes, yourself. Matillion uses a credit-based consumption model tied to the compute it triggers in the destination warehouse. dbt splits between dbt Core, which is open source and free, and dbt Cloud, which is seat-based across Developer, Team, and Enterprise tiers.

Where buyers get it wrong

A frequent mistake is evaluating a single tool against the full job to be done, extract, load, and transform, when most vendors in this category only do part of it well. Airbyte's strength is connector breadth for extract-and-load; pairing it with dbt for transformation is a more common production pattern than expecting either tool to cover the whole pipeline alone.

A second mistake is underweighting schema-evolution handling until a source system changes its schema in production and breaks a pipeline. Airbyte's automatic schema-evolution handling is a specific, checkable capability, worth asking any vendor to demonstrate rather than assume.

A third mistake is choosing a warehouse-pushdown tool like Matillion without confirming it supports the specific destination warehouse in use. Matillion's supported destinations are Snowflake, BigQuery, Redshift, and Delta Lake; a team on a different warehouse needs to confirm compatibility before assuming pushdown transformation will work.

CartographAI's concept explainer on ETL and reverse ETL covers the underlying extract, load, and transform terms in more depth, including the reverse-ETL pattern of moving warehouse data back into operational tools. For related questions about where resolved identity data fits into this stack, see CartographAI's guide to identity resolution.

A few names worth evaluating

Airbyte, Matillion, and dbt Labs are a few names worth evaluating, and the field is larger than this. Fivetran, Census, and RudderStack, covered in CartographAI's ETL explainer, remain active options, and teams already invested in a specific cloud provider often also look at that provider's native data-integration tooling. CartographAI, a free tool agencies and brands use to research vendors across adtech and martech categories, tracks independent assessments across this field for buyers narrowing a shortlist.

FAQ

Do I need both an extract-and-load tool and a transformation tool? In most modern stacks, yes. Airbyte and Fivetran-style tools handle getting data into the warehouse; dbt or a similar transformation layer handles shaping that data into usable models once it's there. Matillion is one of the few tools that covers meaningful ground on both sides, though its connector list is narrower than a specialist like Airbyte.

What's the practical difference between ETL and ELT? ETL transforms data before loading it into the destination; ELT loads data largely as-is and transforms it afterward using the destination's own compute. Most vendors in this category today run ELT, and warehouse compute costs are a real line item to budget for as a result, not just the vendor's own fee.

Is open-source Airbyte a realistic option, or is Cloud required? Self-hosted Airbyte is a usable path and removes licensing cost, but it shifts the operational burden of running and tuning the infrastructure, typically Kubernetes, onto your own team. Teams without infrastructure capacity to spare usually find the cloud product a better trade-off despite the usage-based cost.

Why doesn't dbt show up as a connector-breadth option? Because dbt doesn't extract or load data at all. It operates purely on data already sitting in the warehouse, and teams typically deploy it alongside an extract-and-load tool like Airbyte or Fivetran rather than in place of one.

How should I budget for a credit- or usage-based pricing model? Ask the vendor for typical consumption figures from a customer at your approximate data volume, and separately estimate the warehouse compute cost that pushdown transformation or increased load volume will trigger, since that cost sits outside the vendor's own invoice.

Does schema evolution handling matter if my source systems rarely change? It matters more than expected, because the failure mode is not gradual: a single unannounced schema change from a source system can break a pipeline outright. Confirming how a vendor handles that scenario is worth doing during evaluation rather than after an incident.