SAS to Python Scalable Code Translation April 23,
Description: SAS to Python Scalable Code Translation April 23, 2025 FedCASIC 2025 Reveal Global Consulting Cameron Milne, Ellie Mamantov, John Lynagh www.revealgc.com Agenda www.revealgc.com 2 3 Majority of surveys at the U.S. Census Bureau run in SAS
Related Topics
Download Presentation
"SAS to Python Scalable Code Translation April 23," is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.
Presentation Transcript
slide1. SAS to Python
Scalable Code Translation April 23, 2025 FedCASIC 2025
Reveal Global Consulting
Cameron Milne, Ellie Mamantov, John Lynagh www.revealgc.com<br>
slide2. Agenda www.revealgc.com 2<br>
slide3. 3 Majority of surveys at the U.S. Census Bureau run in SAS (developed in 1960s)
SAS licenses for enterprises are exceedingly expensive
Most universities prioritize Python and R for statistical programming
Census is migrating survey lifecycles to cloud environments Modernizing Systems at the Census Bureau www.revealgc.com<br>
slide4. 4 By-hand…
Rules-based approaches like Abstract Syntax Trees (AST) require hand-crafted rules and are inflexible and time consuming
Neural Machine Translation (NMT) requires massive amounts of parallel data for training which is unavailable for SAS and Python
Large Language Models (LLMs)? Options in Code Translation www.revealgc.com<br>
slide5. www.revealgc.com 5 Sentinel Values (.M, .L, .N, etc.)
Operating over rows vs. vectorized operations
Merging nuances
Dataset size capabilities
Implied logic (e.g. mapping columns to matrices)
Modularized codebases add additional complexity Challenges unique to SAS<br>
slide6. Design Constraints Goals
Outputs are identical
Strong “first-pass” translations
Minimal “last-mile” effort in testing
Preserve logic, comments, and more
Enable LLM to introduce Pythonic concepts 6 Constraints
Locally-hosted models
No feedback loops or access to data
Minimal redevelopment – stakeholders must map existing logic to new code
On-prem server environment (temporarily) www.revealgc.com<br>
slide7. Pipeline Stages 7 Parsing
Rules-Based Conversions
Chunking
Grouping
Prompting
Translation-Documentation
Post-processing www.revealgc.com<br>
slide8. www.revealgc.com 8 Parsing Categories
Single-line comments
Multi-line comments
Steps (data, proc)
Macros
Other code (options, formatting)<br>
slide9. 9 Rules-Based Conversions Controllable areas:
Array instantiations
Dependencies
Don’t give the model a chance to be lazy! www.revealgc.com<br>
slide10. 10 Chunking within Steps Smaller chunks will fit into the context window more easily, but how do you determine where to cut off? ~150 lines ~300 lines ~100 lines ~400 lines ~250 lines ~350 lines www.revealgc.com<br>
slide11. Grouping for Token Optimization After chunking, we want to provide as much context as possible for the model
Grouping strategies:
Fixed token counts
Normalized token counts
Specified patterns (recodes, loops, etc.) 11 www.revealgc.com<br>
slide12. Prompting 12 Key Components
Role Setting
Task Overview
SAS interpretation rules
Python generation guidelines
SAS code
Python vars to remember www.revealgc.com<br>
slide13. 13 Translation-Documentation www.revealgc.com<br>
slide14. Post Processing 14 Post Processing checklist
Move all imports to the top
Insert any leftover documentation
Prune any syntax errors
Apply pep8 formatting www.revealgc.com<br>
slide15. 15 Pipeline Recap www.revealgc.com<br>
slide16. Findings www.revealgc.com 16 SAS Nuances
New concepts require strategy adjustments
Styling differences
Dynamically generated code LLM-Specific
Repetition loops
Instructions ignored/forgotten
Pseudocode
Translation Gaps Workflow
Documentation-led breakdowns
Multi-point failures
Sentinel value plans vary by survey/codebase<br>
slide17. Thank you! www.revealgc.com<br>
Scalable Code Translation April 23, 2025 FedCASIC 2025
Reveal Global Consulting
Cameron Milne, Ellie Mamantov, John Lynagh www.revealgc.com<br>
slide2. Agenda www.revealgc.com 2<br>
slide3. 3 Majority of surveys at the U.S. Census Bureau run in SAS (developed in 1960s)
SAS licenses for enterprises are exceedingly expensive
Most universities prioritize Python and R for statistical programming
Census is migrating survey lifecycles to cloud environments Modernizing Systems at the Census Bureau www.revealgc.com<br>
slide4. 4 By-hand…
Rules-based approaches like Abstract Syntax Trees (AST) require hand-crafted rules and are inflexible and time consuming
Neural Machine Translation (NMT) requires massive amounts of parallel data for training which is unavailable for SAS and Python
Large Language Models (LLMs)? Options in Code Translation www.revealgc.com<br>
slide5. www.revealgc.com 5 Sentinel Values (.M, .L, .N, etc.)
Operating over rows vs. vectorized operations
Merging nuances
Dataset size capabilities
Implied logic (e.g. mapping columns to matrices)
Modularized codebases add additional complexity Challenges unique to SAS<br>
slide6. Design Constraints Goals
Outputs are identical
Strong “first-pass” translations
Minimal “last-mile” effort in testing
Preserve logic, comments, and more
Enable LLM to introduce Pythonic concepts 6 Constraints
Locally-hosted models
No feedback loops or access to data
Minimal redevelopment – stakeholders must map existing logic to new code
On-prem server environment (temporarily) www.revealgc.com<br>
slide7. Pipeline Stages 7 Parsing
Rules-Based Conversions
Chunking
Grouping
Prompting
Translation-Documentation
Post-processing www.revealgc.com<br>
slide8. www.revealgc.com 8 Parsing Categories
Single-line comments
Multi-line comments
Steps (data, proc)
Macros
Other code (options, formatting)<br>
slide9. 9 Rules-Based Conversions Controllable areas:
Array instantiations
Dependencies
Don’t give the model a chance to be lazy! www.revealgc.com<br>
slide10. 10 Chunking within Steps Smaller chunks will fit into the context window more easily, but how do you determine where to cut off? ~150 lines ~300 lines ~100 lines ~400 lines ~250 lines ~350 lines www.revealgc.com<br>
slide11. Grouping for Token Optimization After chunking, we want to provide as much context as possible for the model
Grouping strategies:
Fixed token counts
Normalized token counts
Specified patterns (recodes, loops, etc.) 11 www.revealgc.com<br>
slide12. Prompting 12 Key Components
Role Setting
Task Overview
SAS interpretation rules
Python generation guidelines
SAS code
Python vars to remember www.revealgc.com<br>
slide13. 13 Translation-Documentation www.revealgc.com<br>
slide14. Post Processing 14 Post Processing checklist
Move all imports to the top
Insert any leftover documentation
Prune any syntax errors
Apply pep8 formatting www.revealgc.com<br>
slide15. 15 Pipeline Recap www.revealgc.com<br>
slide16. Findings www.revealgc.com 16 SAS Nuances
New concepts require strategy adjustments
Styling differences
Dynamically generated code LLM-Specific
Repetition loops
Instructions ignored/forgotten
Pseudocode
Translation Gaps Workflow
Documentation-led breakdowns
Multi-point failures
Sentinel value plans vary by survey/codebase<br>
slide17. Thank you! www.revealgc.com<br>