The Library of Congress is searching for a Data Scientist on its Data Analysis and Reporting Team, according to a job posting on USAJobs.gov.
This Data Scientist position is located in the Financial Management Directorate (FMD) in the Library Collections and Services Group (LCSG) and reports to the Director, FMD. The incumbent provides high-level technical expertise to accomplish analysis of data, development of methodological approaches, study design, and advanced written, verbal, and visual communications of study/analysis output.
According to the posting, duties for the Data Scientist include:
- Supports the Enterprise Planning and Management (EPM) initiative to help drive improved planning and management business decisions for the Library of Congress (Library) through advanced data science methods including, but not limited to, anomaly detection, regression or association analysis, record linkage, business analytics, data visualization, predictive analytics (including forecasting), and/or statistical analysis.
- Supports other priority initiatives including, but not limited to, payroll forecasting and modeling, including attrition prediction, as well as cost management and other internal controls.
- Leads or consults with cross-functional teams, staffed by agency staff and contractors, composed of professionals skilled in policy analysis, financial management, information technology, and project management disciplines, in order to develop data-informed solutions that address Library program and business challenges. Generates results from data-driven analyses to produce recommendations to agency management.
- Determines and consults with relevant stakeholders and potential sources to identify the appropriate data, methodological approach, and design. Independently analyzes data, applying and explaining the statistical and mathematical principles used.
- Conducts observational analyses using software and/or programming languages to explore/group data, test hypotheses, predict outcomes, and inform decisions.
- Works with large datasets and numerous confounding and incongruent variables. Derives meaning from data (i.e. datasets that may be large, disparate, unstructured, and/or complex), including structured, loosely structured, and unstructured data).

