Generative AI tools such as Claude Code can accelerate COBOL application analysis, documentation and selected code-generation tasks. Their suitability for translating complete enterprise applications must, however, be evaluated through compilation, execution and functional-equivalence testing.
SoftwareMining is a specialist provider of automated, non-AI, rule-based COBOL-to-Java translation tools, with optional C# generation. We evaluated Claude Code using benchmark programs and representative enterprise COBOL applications containing technologies such as CICS and IMS.
The evaluation compares two different approaches:
Enterprise modernization involves more than generating source code. Small differences in business logic, data handling, transaction processing or batch execution can have significant operational consequences. For most enterprise projects, demonstrating functional equivalence between the original COBOL application and the translated Java or C# application represents a substantial part of the modernization effort.
We evaluated Claude Code using representative enterprise COBOL programs together with the NIST COBOL benchmark suite. The objective was not simply to generate readable Java, but to determine whether the translated applications were complete, compilable, executable and suitable for enterprise modernization.
These results suggest that benchmark results alone is not sufficient to assess enterprise modernization tools. Organizations should evaluate representative application programs, compile and execute the translated system, and compare the results with the original COBOL application.
A meaningful enterprise evaluation should include representative batch and online programs, copybooks, embedded SQL, CICS or IMS transaction processing, representative business data, and execution-based testing to determine functional completeness and the amount of manual correction required.
AI-based tools can support analysis, documentation and developer assistance. Enterprise modernization, however, also requires controlled change management, consistent application architecture and execution-based validation. Before production deployment, the translated application must be tested against the original COBOL system using representative business scenarios and data.
SoftwareMining's automated, non-AI translation tools apply predefined rules consistently across complete COBOL applications. When the same COBOL source, translator version and configuration are used, the tools generate consistent Java output, with optional C# generation.
This repeatability is verified through SoftwareMining's nightly regression process, which translates hundreds of thousands of lines of COBOL test code, including NIST benchmark programs and representative applications. Generated source code and documentation are compared with the previous validated build, while selected applications are compiled, executed and checked for functional correctness and performance regressions. Unexpected differences are flagged for investigation, while intended changes to translation rules and output designs are reviewed before release.
This internal testing supports the stability and repeatability of the translation technology. Each customer application must still undergo project-specific functional-equivalence, integration, security, performance and user-acceptance testing before production deployment.
Every COBOL application is different, and AI-assisted modernization tools continue to evolve. Rather than relying solely on published benchmarks or third-party evaluations, organizations should assess modernization technologies using representative programs from their own applications.
Generated code should be evaluated not only for readability, but also for translation completeness, repeatability, platform coverage and functional equivalence with the original COBOL application.
Publicly available benchmarks, such as the NIST COBOL test suite, can provide an additional test. Removing comments and meaningful identifiers helps determine whether a tool is analyzing program logic or relying on descriptive names and embedded documentation.
Evaluation should also include representative applications containing copybooks, batch and online processing, embedded SQL, CICS or IMS transactions. Whether evaluating SoftwareMining, Claude Code or another modernization technology, the generated Java or C# should be compiled, executed and compared with the original COBOL application to demonstrate complete and functionally equivalent results.
If you're evaluating approaches to COBOL modernization, the following resources may also be useful: