In the early stages of a data platform, the objective is usually pragmatic: connect a source, bring its data into the warehouse and start modelling it. Each data engineer focuses on solving their own integration and builds the process around the specific need in front of them.
This way of working allows rapid progress. However, as the number of sources grows and multiple people develop on the same platform, individually made decisions begin to affect the system as a whole.
Each ingestion process may end up using its own structure, retry policies, error-handling approach and monitoring mechanisms. The processes work and the data reaches its destination, but the platform lacks a common foundation.
While the number of pipelines remains small, these differences can be manageable. The problem emerges as the platform grows: incorporating a new source no longer consists solely of developing its extraction. It also requires deciding again how to start the process, how to coordinate its phases, what information to log and what should happen when something fails.
At that point, individual autonomy can begin to turn into operational complexity.
The problem: growing without a common foundation
This situation arose in a zenital project built on an AWS data platform.
In this environment, Amazon EventBridge scheduled executions, AWS Step Functions coordinated the different processes, and AWS Glue executed the extraction and processing logic. Afterwards, dbt transformed the data, and Power BI used it to refresh and distribute information. Amazon CloudWatch centralised logs and enabled execution monitoring.
As we incorporated new sources, each ingestion process had been built according to its own criteria. All processes fulfilled their purpose, but they could use different structures, apply their own retry policies or handle errors differently.
The lack of a common foundation did not only affect development time. It also made it more difficult to understand what was running at any given moment and increased dependence on the knowledge of the people who had created each process.
When an incident occurred, it was first necessary to understand how that specific ingestion process had been built: what input it received, which jobs it executed, how it propagated errors and which processes depended on it. Two similar failures could require different investigations because the pipelines did not necessarily follow the same pattern.
This variability also made cross-cutting changes more difficult. An improvement in error handling, monitoring or job execution might have to be implemented separately across multiple processes.
As a result, part of the team's effort was spent rebuilding orchestration mechanisms that had already been solved before, rather than focusing on source-specific logic or developing models that delivered business value.
The need, therefore, was not to create a new state machine. It was to establish a shared way of designing and operating ingestion processes.
The solution: standardising without building a monolith
To solve the problem, we evaluated different alternatives.
The first was to maintain an independent state machine for each ingestion process. This approach provided autonomy and allowed each flow to be fully adapted to the characteristics of its source. However, it also preserved duplication and allowed different approaches to continue emerging.
At the opposite extreme, we could build a single fully generic workflow. All sources would pass through the same orchestrator and share a single implementation. This would provide uniformity, but it could also create a monolithic component full of conditions and exceptions.
If every new requirement had to be incorporated into the central flow, it would eventually need to know the details of all sources. Centralisation would reduce some inconsistency, but at the cost of increased coupling, greater change impact and dependency on a single component.
We chose a middle ground: combining a general orchestrator with specialised modules.
The objective was not for all ingestion processes to be identical, but for them to share those operational decisions that made no sense to redefine in every pipeline.
The design principle can be summarised as follows:
Centralise contracts, error management and observability while keeping source-specific logic separate.
How the architecture works
The first step was to identify the capabilities that were repeatedly used across the platform. From these, reusable processes were created for ingestion from databases, APIs and files, as well as for executing dbt jobs, refreshing Power BI and distributing updated reports.
Each module solves a specific responsibility and can be reused across different processes. A new ingestion process does not need to implement the entire chain; it simply combines the capabilities it requires.
The overall flow can be represented as follows:

Amazon EventBridge acts as the scheduling and entry point. In addition to triggering execution, it sends a common contract containing the context required to identify which process is being launched, which sources should run and which phases belong to the flow.
This input reaches a general Step Function, which interprets the requested execution. The engineer defines which modules should be activated and which sources they should operate on. The orchestrator then composes the flow using the available processes.
One execution might, for example, be limited to a file ingestion. Another might combine a relational ingestion, the execution of a dbt job and the subsequent refresh of a report. The overall structure remains the same, but the composition changes according to the requirement.
Extraction logic remains in AWS Glue. Step Functions coordinates the phases and controls the execution state, while each Glue job handles the specifics of its source.
This separation prevents the orchestrator from needing to know how an API authenticates, how a file is processed or which query a relational extraction uses. Its responsibility is to coordinate the components, not absorb their internal logic.
Applied best practices
The official AWS Step Functions documentation recommends dividing complex processes into modular, reusable components with clear responsibilities. We applied this principle to our context through four decisions:
- Distinct responsibilities. Each module represents a recognisable capability, such as relational ingestion, API ingestion or dbt job execution.
- Common interfaces. Modules receive a consistent input structure and preserve the context required to identify the execution.
- Consistent error management. The platform uses a common pattern to detect, propagate and expose failures, although each category of error may require a different response.
- Cross-cutting observability. CloudWatch centralises logs, metrics, execution duration and execution status.
Each module must have a clear responsibility and an understandable interface. Otherwise, fragmentation would simply move complexity from one large workflow into many workflows that are difficult to relate to one another.
Benefits: for the team and for the client
For the team
Before the redesign, adding a source required developing both its specific logic and much of its operational behaviour. The engineer had to decide how to start the process, how to structure the flow, how to execute jobs and how to monitor the outcome.
With the new architecture, a significant portion of these decisions has already been resolved.
The engineer can select the appropriate module for the source type, configure its parameters and implement the extraction's specific logic. When the process requires executing dbt models or refreshing a Power BI report, those modules can be incorporated without having to redevelop their integrations.
This does not mean that all sources become trivial. An API may require a particular authentication mechanism, a database may require a specific incremental strategy and a file may present a complex format.
Those differences still exist, but they no longer require redesigning the entire execution architecture.
For the team, this common foundation provides several benefits:
- Reduces repetitive technical work.
- Simplifies incident investigation.
- Makes it possible to apply improvements to shared components.
- Reduces dependence on the person who originally built each process.
- Maintains autonomy for developing source-specific logic.
Autonomy does not disappear. It becomes focused on the decisions where it genuinely adds value.
For the client
The impact of this architecture is not limited to the technical team.
Adding a new source requires less orchestration work, allowing resources to be dedicated sooner to understanding the data, developing transformations and addressing business needs.
A common structure also improves traceability and makes incident investigation easier. The platform retains the context required to know which process was started, which components were executed and at which phase an error occurred.
For the client, this translates into a platform that is better prepared to grow. New requirements do not require building an architecture from scratch, but rather extending a foundation that already contains the usual operational capabilities.
Conclusion: Governing without eliminating exceptions
This approach also introduces a trade-off. Sharing contracts and processes improves consistency and traceability, but creates a degree of coupling around common components.
This is why the scope of the general orchestrator must remain limited. If it had to be modified every time a special case appeared, it would become the bottleneck we were trying to avoid.
Governance is not about forcing every process to be identical. It is about defining a recommended way of working and ensuring that exceptions are conscious, visible and justified.
In this context, governing ingestion processes means turning the team's criteria into common platform behaviours: how an execution is started, how it is identified, how it is controlled and how it is observed.
The most important change was not creating a new Step Function, but preventing every new source from having to solve the same operational problems again.
A data platform scales when it no longer forces every engineer to make the same operational decisions over and over again.
