Case Study Insurance
Governing Legacy Mainframe COBOL Data for a Tier 1 Health Insurer
Bringing metadata governance, lineage, and automated discovery to decades of legacy mainframe COBOL data feeding a modern Hadoop landing zone
From zero governance to full lineage visibility across every COBOL copybook
Summary
A large, long-standing tier-1 U.S. health insurance provider — an organization generating tens of billions of dollars in annual revenue — had an unconfigured, non-live Informatica Data Catalog and no governance over the legacy mainframe COBOL files feeding a modern HDFS landing zone. Metadata changes were discoverable only through manual documentation, exposing the organization to audit and financial risk. aiDataWorks assessed the environment, gathered stakeholder requirements, then configured Informatica EDC, mainframe, and HDFS connections while building custom COBOL copybook scanners integrated with Informatica Data Processor and EDC. The team wired the solution into the broader Informatica platform and the Control-M corporate scheduler, exposed lineage and metadata in the EDC UI, and supported QA, go-live, and knowledge transfer, followed by post-production managed services. The result was a fully functional, production-live governance environment offering complete lineage and impact visualization of COBOL copybooks and automated metadata-change discovery, with estimated benefits of millions of dollars per year in avoided audit penalties and customer churn.
The challenge
- No governance over legacy mainframe data. Legacy mainframe COBOL files feeding the modern HDFS landing zone had no governance controls in place, leaving a critical data path effectively unmonitored.
- Data catalog not live. The customer's existing Informatica Enterprise Data Catalog installation was unconfigured, had never been smoke-tested, and was not in production use.
- Unsupported metadata scanning. The legacy mainframe file system fell outside standard metadata scanning support, leaving COBOL copybook structures effectively invisible to the catalog.
- Manual metadata tracking. Changes to legacy COBOL metadata were tracked by hand in documentation, creating discovery gaps and audit risk.
The solution
How we built it
Stakeholder assessment & requirements
- Interviewed major stakeholders to understand prior work on the catalog and current expectations for the governance program.
- Performed a technical assessment of the existing Informatica and mainframe environment.
- Produced a consolidated requirements and assessment document, reviewed and validated with stakeholders before implementation began.
EDC & mainframe connectivity setup
- Configured and smoke-tested Informatica Enterprise Data Catalog (EDC) 10.1 from the ground up.
- Established and validated the mainframe source connection against the customer's COBOL file systems.
- Configured the HDFS landing zone targets on the Big Data Management/Big Data Engineering (BDE) 10.1 environment, fixing issues as they were found.
Custom COBOL copybook scanner build
- Developed and implemented custom COBOL copybook scanners using the customer's existing code together with Informatica Data Processor (Data Transformation).
- Integrated the scanners with EDC so copybook structures could be catalogued and searched like any other governed asset.
- Loaded scan results into the target EDC HDFS landing zone for downstream lineage and impact analysis.
Platform integration, QA & managed services rollout
- Integrated the solution with the broader Informatica platform and the corporate scheduler, Control-M.
- Exposed scan results, lineage, and copybook relationships in the EDC user interface.
- Performed QA, supported go-live, delivered knowledge transfer to the customer's team, and provided post-production managed services.
Outcomes
- The governance environment is now fully functional and production-live, replacing an Informatica Data Catalog installation that had never been configured or smoke-tested.
- Any number of COBOL copybooks can now be visualized for lineage, impact, and relationships — including nested records, data types, dependencies, and maintenance logs — where none of that visibility existed before.
- Metadata changes are now discovered automatically through scheduled EDC/Informatica workflow scans rather than manual documentation, with the customer estimating the shift avoids millions of dollars per year in audit penalties and customer churn.
Solving something similar?
We will walk your current data landscape and show you what a governed customer view would take.
