- Haiyan Meng presented Conducting Reproducible Research with Umbrella: Tracking, Creating, and Preserving Execution Environments, which describes how the Umbrella framework is used to create a precise specification of computational environment, software, and data for reproducible execution for epidemiology and high energy physics codes.
- Peter Ivie presented PRUNE: A Preserving Run Environment for Reproducible Computing which describes PRUNE, a workflow system that tracks both data and executions in a way that can be compactly named and shared. This allows one to uniquely identify an execution in a way that others can track the complete provenance, or re-execute it if desired.
Tuesday, October 25, 2016
Reproducibility Papers at eScience 2016
CCL students presented two papers at the IEEE 12th International Conference on eScience on the theme of reproducibility in computational science:
CCL Workshop 2016
The 2016 CCL Workshop on Scalable Scientific Computing was held on October 19-20 at the University of Notre Dame. We offered tutorials on Makeflow, Work Queue, and Parrot. and gave highlights of the many new capabilities relating to reproducibility and container technologies. Our user community gave presentations describing how these technologies are used to accelerate discovery in genomics, high energy physics, molecular dynamics, and more. Everyone got together to share a meal, solve problems, and generate new ideas. Thanks to everyone who participated, and see you next year!
Tuesday, September 20, 2016
NSF Grant to Support CCTools Development
We are pleased to announce that our work will continue to be supported by the National Science Foundation through the division of Advanced Cyber Infrastructure.
The project is titled "SI2-SSE: Scaling up Science on Cyberinfrastructure with the Cooperative Computing Tools" It will advance the development of the Cooperative Computing Tools to meet the changing technology landscape in three key respects: exploiting container technologies, making efficient use of local concurrency, and performing capacity management at the workflow scale. We will continue to focus on active user communities in high energy physics, which rely on Parrot for global scale filesystem access in campus clusters and the Open Science Grid; bioinformatics users executing complex workflows via the VectorBase, LifeMapper, and CyVerse disciplinary portals, and ensemble molecular dynamics applications that harness GPUs from XSEDE and commercial clouds.
The project is titled "SI2-SSE: Scaling up Science on Cyberinfrastructure with the Cooperative Computing Tools" It will advance the development of the Cooperative Computing Tools to meet the changing technology landscape in three key respects: exploiting container technologies, making efficient use of local concurrency, and performing capacity management at the workflow scale. We will continue to focus on active user communities in high energy physics, which rely on Parrot for global scale filesystem access in campus clusters and the Open Science Grid; bioinformatics users executing complex workflows via the VectorBase, LifeMapper, and CyVerse disciplinary portals, and ensemble molecular dynamics applications that harness GPUs from XSEDE and commercial clouds.
Thursday, September 15, 2016
Announcement: CCTools 6.0.0. released
The Cooperative Computing Lab is pleased to announce the release of version 6.0.0 of the Cooperative Computing Tools including Parrot, Chirp, Makeflow, WorkQueue, Umbrella, Prune, SAND, All-Pairs, Weaver, and other software.
The software may be downloaded here:
http://ccl.cse.nd.edu/software/download
This is a major which adds several features and bug fixes. Among them:
We will have tutorials on the new features in our upcoming workshop, October 19 and 20. Refer to http://ccl.cse.nd.edu/workshop/2016 for more information. We hope you can join us!
Thanks goes to the contributors for many features, bug fixes, and tests:
Please send any feedback to the CCTools discussion mailing list:
http://ccl.cse.nd.edu/community/forum
Enjoy!
The software may be downloaded here:
http://ccl.cse.nd.edu/software/download
This is a major which adds several features and bug fixes. Among them:
- [Catalog] Automatic fallback to a backup catalog server. (Tim Shaffer)
- [Makeflow] Accept DAGs in JSON format. (Tim Shaffer)
- [Makeflow] Multiple documentation omission bugs. (Nick Hazekamp and Haiyan Meng)
- [Makeflow] Send information to catalog server. (Kyle Sweeney)
- [Makeflow] Syntax directives (e.g. .SIZE for to indicate file size). (Nick Hazekamp)
- [Parrot] Fix cvmfs logging redirection. (Jakob Blomer)
- [Parrot] Multiple bug-fixes. (Tim Shaffer, Patrick Donnelly, Douglas Thain)
- [Parrot] Timewarp mode for reproducible runs. (Douglas Thain)
- [Parrot] Use new libcvmfs interfaces if available. (Jakob Blomer)
- [Prune] Use SQLite as backend. (Peter Ivie)
- [Resource Monitor] Record the time where a resource peak occurs. (Ben Tovar)
- [Resource Monitor] Report the peak number of cores used. (Ben Tovar)
- [Work Queue] Add a transactions log. (Ben Tovar)
- [Work Queue] Automatic resource labeling and monitoring. (Ben Tovar)
- [Work Queue] Better capacity worker autoregulation. (Ben Tovar)
- [Work Queue] Creation of disk allocation per tasks. (Nate Herman-Kremer)
- [Work Queue] Extensive updates to wq_maker. (Nick Hazekamp)
- [Work Queue] Improvements in computing master's task capacity. (Nate Herman-Kremer).
- [Work Queue] Raspberry Pi compilation fixes. (Peter Bui)
- [Work Queue] Throttle work_queue_factory with --workers-per-cycle. (Ben Tovar)
- [Work Queue] Unlabeled tasks are assumed to consume 1 core, 512 MB RAM and 512 MB disk. (Ben Tovar)
- [Work Queue] Worker disconnects when node does not longer have the resources promised. (Ben Tovar)
- [Work Queue] work queue statistics clean up (see work_queue.h for deprecated names). (Ben Tovar)
- [Work Queue] work_queue_status respects terminal column settings. (Mathias Wolf)
We will have tutorials on the new features in our upcoming workshop, October 19 and 20. Refer to http://ccl.cse.nd.edu/workshop/2016 for more information. We hope you can join us!
Thanks goes to the contributors for many features, bug fixes, and tests:
- Jakob Blomer
- Peter Bui
- Patrick Donnelly
- Nathaniel Kremer-Herman
- Kenyi Hurtado-Anampa
- Peter Ivie
- Kevin Lannon
- Haiyan Meng
- Tim Shaffer
- Douglas Thain
- Ben Tovar
- Kyle Sweeney
- Mathias Wolf
- Anna Woodard
- Chao Zheng
Please send any feedback to the CCTools discussion mailing list:
http://ccl.cse.nd.edu/community/forum
Enjoy!
Friday, August 19, 2016
Summer REU Projects in Data Intensive Scientific Computing
We recently wrapped up the first edition of the summer REU in Data Intensive Scientific Computing at the University of Notre Dame. Ten undergraduate students came to ND from around the country and worked on projects encompassing physics, astronomy, bioinformatics, network sciences, molecular dynamics, and data visualization with faculty at Notre Dame.
To learn more, see these videos and posters produced by the students:
To learn more, see these videos and posters produced by the students:
Wednesday, August 10, 2016
Simulation of HP24stab with AWE and Work Queue
The villin headpiece subdomain "HP24stab" is a recently discovered 24-residue stable
supersecondary structure that consists of two helices joined by a turn.
Simulating 1μs of motion for HP24stab can take days or weeks depending on the
available hardware, and folding events take place on a scale of hundreds of
nanoseconds to microseconds. Using the Accelerated Weighted Ensemble (AWE), a total of 19us of
trajectory data were simulated over the course of two months using the OpenMM simulation
package. These trajectories were then clustered and sampled to create an AWE
system of 1000 states and 10 models per state. A Work Queue master dispatched 10,000 simulations to a peak of 1000 connected 4-core workers, for a total of 250ns of
concurrent simulation time and 2.5μs per AWE iteration. As of August 8, 2016,
the system has run continuously for 18 days and completed 71 iterations, for a
total of 177.5μs of simulation time. The data gathered from these simulations
will be used to report the mean first passage time, or average time to fold,
for HP24stab, as well as the major folding pathways. - Jeff Kinnison and Jesús
Izaguirre, University of Notre Dame
ND Leads DOE Grant on Virtual Clusters for Scientific Computing
Our current NSF and DOE supercomputers are very powerful, but they each have different operating systems and software configurations, which makes it difficult and time consuming for new users to deploy their codes and share results. The new service will create virtual clusters on the existing machines that have the custom software and other services needed to easily run advanced scientific codes from fields such as high energy physics, bioinformatics, and astrophysics. If successful, users of this service will be able to easily move applications between university and national supercomputing facilities.
Friday, July 29, 2016
2016 DISC Summer Session Wraps Up
At our closing poster session in Jordan hall (along with several other REU programs) students presented their work and results to faculty and guests across campus.
If you are excited to work at the intersection of scientific research and advanced computing, we invite you to apply to the 2017 DISC summer program at Notre Dame!
Wednesday, June 22, 2016
New Work Queue Visualization
Nate Kremer-Herman has created a new, convenient way to lookup information of Work Queue masters. This new visualization tool provides real-time updates on the status of each Work Queue master that contacts our catalog server. We hope that this new tool will serve to both facilitate our users' understanding of what their Work Queue masters are doing and assist the user in determining when it may be time to take corrective action.
In part, this tool provides our users with measurements on their tasks currently running, the number of tasks waiting to be run, and the total capacity of tasks that could be running. As an example, a user could find that they have a large number of tasks waiting, a small number of tasks running, and a task capacity that is somewhere in between. A recommendation we could make to a user who is seeing something like this would be to ask for more workers. Our hope is that users will take advantage of this new way to view and manage their work.
![]() |
| A Comparative View |
![]() |
| A Specific Master |
In part, this tool provides our users with measurements on their tasks currently running, the number of tasks waiting to be run, and the total capacity of tasks that could be running. As an example, a user could find that they have a large number of tasks waiting, a small number of tasks running, and a task capacity that is somewhere in between. A recommendation we could make to a user who is seeing something like this would be to ask for more workers. Our hope is that users will take advantage of this new way to view and manage their work.
Thursday, May 26, 2016
Work Queue from Raspberry Pi to Azure at SPU
"At
Seattle Pacific University we have used Work Queue in the CSC/CPE 4760
Advanced Computer Architecture course in Spring 2014 and Spring 2016.
Work Queue serves
as our primary example of a distributed system in our “Distributed and
Cloud Computing” unit for the course. Work Queue was chosen because it
is easy to deploy, and undergraduate students can quickly get started
working on projects that harness the power
of distributed resources."
The main project in this unit had the students obtain benchmark results for three systems: a high performance workstation; a cluster of 12 Raspberry Pi 2 boards, and a cluster of A1 instances in Microsoft Azure. The task for each benchmark used Dr. Peter Bui’s Work Queue MapReduce framework; the students tested both a Word Count and Inverted Index on the Linux kernel source. In testing the three systems the students were exposed to the principles of distributed computing and the MapReduce model as they investigated tradeoffs in price, performance, and overhead.
- Prof. Aaron Dingler, Seattle Pacific University.
The main project in this unit had the students obtain benchmark results for three systems: a high performance workstation; a cluster of 12 Raspberry Pi 2 boards, and a cluster of A1 instances in Microsoft Azure. The task for each benchmark used Dr. Peter Bui’s Work Queue MapReduce framework; the students tested both a Word Count and Inverted Index on the Linux kernel source. In testing the three systems the students were exposed to the principles of distributed computing and the MapReduce model as they investigated tradeoffs in price, performance, and overhead.
- Prof. Aaron Dingler, Seattle Pacific University.
Subscribe to:
Posts (Atom)







