This CSV file contains the topic distribution of each EIN as uncovered using six parallel Latent Dirichlet Allocation (LDA) Topic Models.
Each row depicts a topic and topic-score associated with an Ohio NPO (identified by Employer Identification Number) generated from one model run.
The sum of topic scores possible for every row associated with an EIN therefore will not exceed 6.0 (6 models x 100%)
Topic scores below .01 (1%) are not included.
Each topic from the models is further identified as Essential/Non-Essential by subject matter expert, Dr. Michael Jones, guided by the official IRS definition.
The topic models are generated on unstructured text language from the mission statement and activities language taken from the 2019 tax forms of Ohio non-profit organizations.
This webinar was a part of the Data and Computation Science Series. It occurred on March 4, 2021, at 2:00 pm EST.
Presenter Bio for Ashley Farley:
Over the past decade, Ashley has worked in both academic and public libraries, focusing on digital inclusion and facilitating access to scholarly content. She completed her Masters's in Library and Information Sciences through the University of Washington’s Information School.
Ashley is a Program Officer of Knowledge and Research Services at the Bill and Melinda Gates Foundation. In this capacity, she leads the foundation’s Open Access Policy’s implementation and associated initiatives. This includes leading the work of Gates Open Research, a transparent and revolutionary publishing platform. Other core activities involve supporting the strategic and operational aspects of the foundation’s library. This work has sparked a passion for open access, believing that freely accessible knowledge has the power to improve and save lives.”
Title of Presentation: Open Research: Making Harmful Habits History
All models and corresponding network visualizations are generated from documents in the CORD-19 dataset as of July 14, 2020. All annotations in red were added by the research team.
Note: These topic models are included here as additional reference and to append links to interactive versions on the Digital Scholarship Center’s machine learning platform for further exploration.
All models and corresponding network visualizations are generated from virus related documents in the CORD-19 dataset as of July 14, 2020. All annotations in red were added by the research team.
Note: Certain Non-Coronaviridae topic models are included in the text of this article and are included here only as additional reference and to append links to interactive versions on the Digital Scholarship Center’s machine learning platform for further exploration.
All models and corresponding network visualizations are generated from virus related documents in the CORD-19 dataset as of July, 2020. All annotations in red were added by the research team.
Note: Coronavirus topic models are included in the text of this article and are included here only as additional reference and to append links to interactive versions on the Digital Scholarship Center’s machine learning platform for further exploration.
These Centrality measurements were generated with NetworkX, a Python package for networks. The specific algorithms used for this paper are Betweenness Centrality (where Degree Centrality considers individual topics).
Complete Centrality Data for this research can be found at https://scholar.uc.edu/show/6t053h21x
A presentation from the Society of Ohio Archivists 2020 meeting.
The University of Akron University Archives and the University of Cincinnati Libraries will present and analyze challenges faced by institutions looking to create, implement and improve their digital preservation program. Armed with the NDSA Levels of Digital Preservation and the Digital Preservation Capability Maturity Model (DPCMM), both institutions discuss strategies to tackle common issues such as minimal staffing, limited resources, procrastination, and legacy digital content. Each institution will also discuss strategies used to handle unique challenges faced in crafting their individual digital preservation policies.
Presentation recording available: https://youtu.be/czemLLqXNh8
A presentation at the joint Upper Midwest Digital Collections Conference and Minnesota Digital Library Annual Meeting in 2020.
The diversity of a digital collection is often assessed by considering the diversity of its content. In order for collections to be truly inclusive, however, they need to emphasize usability alongside broad representation. The University of Cincinnati Libraries discusses how diversity and accessibility are intersectional considerations of digital collections, and introduces newly implemented workflows and standards designed to create accessible, inclusive digital collections that broaden usability for all.
Presentation recording available (Starts at 14 minutes, 30 seconds): https://youtu.be/srIPaD7RvYo
The files in this work represent the presentations and workshop content from the 5th UC Data Day held 2020-10-23.
The theme was “World Changing Data: How Digital Data Will Change Our Future”.
The Keynote speaker was Glenn Ricart, of US Ignite - "Smart Runs on Data"
Interactive Panel featuring: Michael Dunaway (moderator) - Whitney Gaskins (Asst Dean, CEAS - Incl Excellence & Comm Engagmnt) - Zvi Biener (Assoc Professor, A&S Philosophy) - Prashant Khare (Asst Professor, CEAS - Aerospace Eng & Eng Mechanics)- Sam Anand (Professor, CEAS - Mechanical Eng) - Achala Vagal
(Professor Clinical - GEO, COM Radiology Neuroradiology)
Power Sessions:
George Turner - Indiana University - High-Performance Computing at UC
Erin McCabe - University of Cincinnati - Text Mining, Natural Language Processing & AI
link to slides - https://bit.ly/dataday_slides
link to code - https://bit.ly/dataday_code
Videos of the day can be found on the UC Libraries STRC1 youtube channel - https://www.youtube.com/c/STRC1/videos