When native cloud schedulers are enough, and when not
A five-point test for deciding whether the scheduler that ships with your platform is enough.
Every cloud platform ships with a scheduler, and for work that stays inside that platform it is usually the right tool. The trouble begins at the boundary, where one platform's scheduler has to wait on another platform's work and no single graph covers the whole run.
Start With What the Native Scheduler Gets Right
It is worth saying plainly: a scheduler that ships with the platform is not a compromise. There is no extra service to run, authentication and monitoring already understand the platform's own resources, retries happen where the failures happen, and the platform's team keeps it working.
So the useful question is never "is the native scheduler good?" It is "how much of my workflow stays inside one boundary?"
The Boundary Test
Five checks. If all five hold, the native scheduler is enough. Each one that fails is a place where you are about to start building glue.
You are still comfortably inside one platform when:
- every step is a first-class job in the same platform, not an HTTP call;
- a single dependency graph covers the whole run;
- monitoring shows the whole run, not one platform's slice of it;
- retries and idempotency are handled inside the platform;
- adding a step does not mean inventing a new integration pattern.
You are past the edge when:
- steps are chained with storage events, webhooks or polling timers;
- a step is a service call rather than a job the scheduler understands;
- "did the other platform's job finish?" is answered by a person, or by a table someone maintains;
- lineage and run history stop at the platform boundary;
- adding the next platform means building another bespoke bridge.
Scenario 1: One Platform Does the Whole Job
Picture a pipeline where every stage lives in Databricks. Data is ingested from cloud storage and message buses by the platform's declarative pipelines. Transformations run in the same pipelines. Unstructured documents are parsed with an AI function, ai_parse_document. The parsed output is indexed for search with Databricks AI Search. Analysts read the result in AI/BI dashboards.
Lakeflow Jobs schedules the lot with one dependency graph, and the platform's monitoring shows every task in one place. Run the five checks: all five hold.
In this shape, adding an external orchestrator removes nothing and adds a hop, a credential and a failure mode. "Just use the native scheduler" is the correct answer, and it is worth saying so.
Scenario 2: The Same Pipeline, Two Platforms
Now mirror that pipeline onto a mixed architecture. Ingestion runs in Azure Data Factory, copying source data into a storage account. Document extraction runs on Azure Document Intelligence. Transformation runs on Databricks. The search index is built in Azure AI Search. Dashboards live in Power BI.
Logically it is the same pipeline as scenario 1. Physically it crosses three or four boundaries, and that changes everything.
Data Factory can invoke a Databricks job directly, and it has a connector that writes to an Azure AI Search index, so those hops are expressible inside the platform. Document extraction has no first-class activity, so it becomes a Web call: a place where the dependency no longer resolves on its own. Deeper in, the Databricks job schedules its own tasks on its own clock, and Power BI refreshes on a third.
Run the five checks again. The graph no longer covers the run. Monitoring is split across four consoles. "Did extraction finish?" needs a lookup. The next service you add needs another bridge. None of this is a defect in any one product. It is what a boundary does.
One platform
- Every stage is a job in the same platform
- One dependency graph covers the run
- One place shows every task
- Retries handled by the platform
- Adding a step needs no new pattern
Mixed platforms
- Steps chained by events, webhooks or timers
- The graph stops at each boundary
- Four consoles to see one run
- "Did the other job finish?" needs a lookup
- Each new platform needs another bridge
Three Ways to Cross a Boundary
- Event glue. Storage events, webhooks or a polling timer pick up where the last platform left off. Cheapest to start, and the cost shows up later as silent gaps and manual reconciliation.
- A scheduler of schedulers. One job whose only task is to start other schedulers and wait. It works, but each platform still sees only its own slice, and the outer graph is a script that one person owns.
- A control plane that spans the platforms. Each platform stays the executor. A separate layer holds the dependency graph across them, maps parameters between tasks, and keeps one run history.
Choosing Between Them
Use event glue when the boundary is a single hop and happens rarely. Use a scheduler of schedulers only if you accept that its view is partial. Reach for a cross-platform control plane when the graph matters, that is, when "run this only after that finished on another platform" is a requirement rather than a convenience.
A simple test: try to answer, for one run, what ran, in what order, and what waited on what. If the honest answer needs four browser tabs, the boundary is costing more than an orchestrator would.
Where Polysync Fits
Polysync is a control plane rather than a replacement for the platforms you already use. It treats pipelines, jobs and functions across Azure, Databricks, Fabric, AWS and Google Cloud as first-class nodes in one dependency graph, maps parameters between them, applies concurrency budgets across platforms, and shows every run in one place. It does not move your data and does not change how any platform executes its work. It decides what runs, when, and what waits. The platforms you trust stay exactly where they are.
Related: ADF vs Polysync: when ADF stops being enough · What is ETL orchestration?
Microsoft, Azure, Azure AI Search, Azure Data Factory, Azure Document Intelligence, Microsoft Fabric and Power BI are trademarks of the Microsoft group of companies. All other product names, logos and brands are the property of their respective owners and are used for identification purposes only.