The Oracle profiler PR (#2187) introduced a pattern where every logical extractor is split into two YAML steps that share the same name: — one with type: ddl, one with type: sql. This works today because pipeline.py dispatches on step.type first, but it causes three problems:
- Ambiguous failure reporting.
PipelineClass.execute() aggregates StepExecutionResult by step.name, so when the SQL half of config_containers fails, users see Pipeline execution failed due to errors in steps: config_containers with no signal whether DDL or SQL was the culprit.
- Implicit name=table contract.
_execute_ddl_step checks if not self._table_exists(conn, step.name): conn.execute(ddl), so the DDL must CREATE TABLE whose name is exactly step.name. Nothing enforces this; an edit that changes the name in one place but not the other silently no-ops or fails at insert time.
- Future fragility. Anyone refactoring
pipeline.py to dedupe steps by name (perfectly reasonable assumption) breaks the YAML silently.
Scope / acceptance criteria
References
The Oracle profiler PR (#2187) introduced a pattern where every logical extractor is split into two YAML steps that share the same
name:— one withtype: ddl, one withtype: sql. This works today becausepipeline.pydispatches onstep.typefirst, but it causes three problems:PipelineClass.execute()aggregatesStepExecutionResultbystep.name, so when the SQL half ofconfig_containersfails, users seePipeline execution failed due to errors in steps: config_containerswith no signal whether DDL or SQL was the culprit._execute_ddl_stepchecksif not self._table_exists(conn, step.name): conn.execute(ddl), so the DDL mustCREATE TABLEwhose name is exactlystep.name. Nothing enforces this; an edit that changes the name in one place but not the other silently no-ops or fails at insert time.pipeline.pyto dedupe steps byname(perfectly reasonable assumption) breaks the YAML silently.Scope / acceptance criteria
setup_ddl:andextract_sql:fields. Larger refactor, but conceptually cleaner since DDL+SQL are one extractor.type: sql(see Convert MSSQL profiler from type:python venv steps to type:sql in-process steps #2459).References
oracle/pipeline_config.yml:11.type: pythontotype: sql).