Data Engineering Trends for 2027: 5 Shifts Already Underway

5 min read

Trend lists for the coming year usually start from forecasts. This one starts from what shipped in the last 12 months. In data engineering, most of 2027 is already announced: the agents are shipping, and the deadlines are written into law.

The 5 shifts below are all moving now, and several have a date coming up that will make them harder to ignore.

 

1. The platforms ship agents that write pipelines

Snowflake launched Cortex Code in November 2025. In February, it added support for dbt and Airflow, and a subscription for teams that do not run on Snowflake. Google's BigQuery has a data engineering agent that builds pipelines and troubleshoots them. In June, Databricks announced an agent that traces a pipeline failure to its cause and proposes a fix, tested in a sandbox before a person applies it. In September, dbt Labs announced dbt Wizard, and at the end of the month Microsoft put its data engineering agent for Fabric, Project Osmos, into preview.

The interesting part is what these agents now do after writing the code. Microsoft tells users to put their validation criteria in the request: row counts, null checks, uniqueness, reconciliation or business rules. In its own worked example, the agent runs the old and new pipelines against staged targets and compares their row counts, duplicate keys, schema compatibility and output totals. dbt Wizard checks the upstream and downstream impact and builds the change before anyone sees the diff. Verification is moving into the agent itself.

In September, Gartner forecast that by 2029, AI agents will have automated 75% of data engineering workflows. The next wave of announcements is already scheduled, with Microsoft Ignite from 17 November and AWS re:Invent from 30 November.

What none of the launches settles is ownership. Once an agent has written the pipeline and its checks, someone still has to accept code that nobody on the team wrote, and the engineer's job moves from directing an assistant to accepting what it delivers.

 

2. The connector stops being the product

For most of the last decade, the connector was what data integration vendors sold. They maintained it, and you often paid by the volume of rows that went through it.

That is changing from 2 directions. Fivetran's AI Connector Agent, in beta since May, generates new connectors on request, which then run on Fivetran's infrastructure. Databricks opened community connectors in beta: they are built with AI-assisted workflows for Claude Code or Cursor, tested against live sources, and maintained by the community without a Databricks SLA.

When a long-tail connector can be generated in an afternoon, the scarce part is everything around it: the tests, the documentation, the decisions about what the data means, and someone who answers for it when the source changes. The cost of a connector moves from the licence to the maintenance.

The question to ask of each tool is what it actually hands over. A connector hosted by a vendor and a connector delivered as code in your repository look the same in a demo, and they are different families of tool the day the source changes. In 2027, the build-or-buy decision will be made connector by connector.

 

3. The stack consolidates

In less than a year, the data stack lost several of its independent vendors. Salesforce closed its acquisition of Informatica in November 2025. In January 2026, Microsoft bought Osmos, a startup that used AI agents to automate data engineering. IBM completed its $11 billion purchase of Confluent in March, to feed its platform with streaming data for AI. On 1 June, Fivetran and dbt Labs completed their merger, putting 2 of the best-known tools in the stack, one for ingestion and one for transformation, in the same company. And in July, SAP completed its acquisition of Dremio, an open data lakehouse platform.

The picture that emerges is a handful of platforms, each with an agent on top, covering ingestion, transformation, streaming and governance in one contract. For a data team, that means fewer integration seams to maintain, and more dependence on a single vendor's roadmap and pricing.

The independents that remain will increasingly be judged on 2 things: how well they fit inside one of these platforms, and how easily you can leave them. The second question is where regulation comes in.

 

4. Portability moves into law

Connected products placed on the EU market after 12 September 2026 must be designed so that users can access the data they generate. From 12 January 2027, the Data Act stops cloud providers from charging customers for switching to another provider, data egress included. In financial services, DORA already requires exit plans for critical ICT services, and they must be "sufficiently tested".

Open table formats make the storage side of leaving easier. Iceberg v3, approved by the Apache Iceberg community in 2025, brings deletion vectors and row lineage that line up with Delta Lake's, and Snowflake made v3 generally available in May. Data held in an open format can be read by another engine without a migration project.

Storage is the easy part of leaving. The pipelines are harder, because vendor lock-in is built one connector at a time: connectors that only run on one platform, and checks that live in its interface. In 2027, the ability to leave becomes something an auditor or a regulator can ask you to demonstrate.

 

5. Context becomes something data teams ship

Agents need to know what the data means, and over the last year the industry started treating that meaning as something to build and deliver.

In January, the Open Semantic Interchange working group, started by Snowflake and now including Salesforce, dbt Labs and Databricks, published the first version of its specification, a vendor-neutral model for semantic layer definitions, under an Apache 2 licence. In September, Fivetran previewed a Context Layer that assembles context for agents from semantic layers, documents and connected applications. dbt's MCP server, now generally available, exposes a project's models, metrics, lineage and test results to AI agents. Gartner expects that by 2030, universal semantic layers will be treated as critical infrastructure, alongside data platforms and cybersecurity.

For data teams, the deliverable grows. Tables and dashboards are no longer enough. The definitions behind them, and the lineage that explains each value, have to be readable by a machine, for the same reason a code-generating agent needs a harness that knows the domain.

There is a limit worth keeping in mind as these layers arrive. Context added above the warehouse cannot restore what a pipeline discarded on the way in: the rows it dropped, the types it forced, the keys it invented, the nulls that carry no reason. In 2027, the quality of an agent's answers will depend on decisions made at ingestion.

 

What the 5 have in common

Read together, the 5 trends point the same way. Building pipelines gets cheaper every quarter. What stays expensive is ownership: of the code an agent writes, of the checks that say it is right, of the right to leave with both, and of the meaning of the data.

4 questions are worth asking before 2027:

  • When an agent writes a pipeline for you, where does the code live, and can your team read and change it?
  • Who writes the acceptance criteria, and do they live with the code or inside a vendor's interface? The answer settles what production-ready means for your team, and who gets to decide.
  • If one of your pipelines had to move to another platform next quarter, what would you have to rewrite?
  • Which of your definitions and lineage could an agent read today?

 


PrettyWhale.AI is an Engineering AI for ingestion code. Given a sample of the source, it generates the pipeline with its tests and documentation in your repository, so the code and its checks stay with your team, whatever platform runs them.

 

You have any questions on PrettyWhale.ai ?

Contact us