Showing posts with label science. Show all posts
Showing posts with label science. Show all posts

Friday, August 3, 2012

7 Habits of the Open Scientist: #2 -- Reproducible Research

Note: this post is part of a series on habits of the open scientist.  Here I discuss the second habit, reproducible research.  The previous post was on open scientific publishing.

Reproducible research

Reproducibility is part of the definition of science: if the results of your experiment cannot be replicated by different people in a different location, then you're not doing science.  Far from being a mere philosophic concern, reproducible research has been a key issue in prominent controversies like climategate and cancer research clinical trials.

Especially disconcerting is the typical irreproducibility of scientific work involving computer code:

“Computational science is facing a credibility crisis: it’s impossible to verify most of the computational results presented at conferences and in papers today.” (LeVeque, Mitchell, Stodden, CiSE 2012)

Frankly, I used to find that I was often unable to reproduce my own computational results after a few months, because I had not maintained sufficiently detailed notes about my code and my computing environment.

The open scientist ensures that the entire research compendium -- including not only the paper but the data, source code, parameters, post-processing, and computing environment -- is made freely available, preferably in a way that facilitates its reuse by others.

I won't spend more time motivating reproducible research, since others have done that much better than I could.  Instead, let me focus on the relatively easy first steps you can take to make your research more reproducible.

The bare minimum: publish your code and data

If you wish to set an example of good reproducible computational research practices, I have good news for you: the bar is very low at the moment.  The reason why "it's impossible to verify most of the computational results" is that most researchers don't release their code and data.  The first step toward working reproducibly is simply to put the code and data that is used in your published research out in the open.

If you don't want to release your code to the public, please read about why you should and why you can.  Once you're convinced, go endorse the Science Code Manifesto.

Releasing your code and data can be as simple as posting a tarball on your website with a reference to the paper it pertains to.  Or you may wish to start putting all your code out in the open on Bitbucket or Github, like I do.  I don't claim that these are the best solutions possible, but they are a big step forward from keeping everything on your own hard drive.

When you release your code and data, it is important to use an appropriate license.  Victoria Stodden, a leader in the reproducible research movement, recommends the use of a permissive license like modified BSD for code and Science Commons Database Protocol for data.  Together with the Creative Commons BY license for media (that I mentioned in my last post), these comprise the Reproducible Research Standard, a convenient amalgamation of licenses for open science.

Be sure to include a mention of reproducibility in your paper, along with links to the code and data.  If you release your work under the RRS, I suggest using this citation.

Real benefits

The open scientist may adopt reproducible research practices for philosophical reasons, but he soon finds that they bring more direct benefits.  Because he writes code and prepares data with the expectation that it will be seen by others, the open scientist finds it much easier for himself, students, and colleagues to build on past work.  New collaborations are formed when others discover his work through openly released code and data.  And (as in the case of this paper, for example) the code itself may be the main subject of publications in journals that have come to recognize the importance of scientific software.

Taking it further

Like free and open scientific publishing, reproducible research has become a very large movement, and only a book could hope to cover it all.  Here I've merely distilled some basic practical suggestions.

Openly releasing code and data is only the first step.  Open scientists may wish to adopt tools that track code provenance and ensure a fully reproducible workflow, such as

Tuesday, July 31, 2012

7 Habits of the Open Scientist

Science has always been based on a fundamental culture of openness.  The scientific community rewards individuals for sharing their discoveries through perpetual attribution, and the community benefits by through the ability to build on discoveries made by individuals.  Furthermore, scientific discoveries are not generally accepted until they have been verified or reproduced independently, which requires open communication.

 

Historically, openness simply meant publishing one's methods and results in the scientific literature.  This enabled scientists all over the world to learn about essential advances made by their colleagues, modulo a few barriers.  One needed to have access to expensive library collections, to spend substantial time and effort searching the literature, and to wait while research conducted by other groups was refereed, published, and distributed.

 

Nowadays it is possible to practice a fundamentally more open kind of research -- one in which we have immediate, free, indexed, universal access to scientific discoveries.  The new vision of open science is painted in lucid tones in Michael Nielsen's Reinventing Discovery.  After reading Nielsen's book, I was hungry to begin practicing open science, but not exactly sure where to start.  Here are seven ways I'm aware of.  Each will be the subject of a longer forthcoming post.

 

I believe that every scientist has a moral imperative to adopt the first two:

 

1. Freely accessible publications.  At a minimum, make sure that everyone is allowed to read your research.

2. Reproducible research.  Release your code and data so  that anyone who wants to can verify or build directly on your work.

 

The remaining five are marks of a truly open scientist:

 

3. Pre-publication dissemination of research.  Just because peer-review and journals take time, that doesn't mean you need to embargo your audience.

4. Open collaboration through social media.  Find the person who knows that one thing you need, through new scientific networking tools -- and share your own expertise where it's needed most.

5. Live open science.  Tell people about your marvelous discoveries -- as you make them.

6. Open expository writing.  Teach others about the field you work in through a blog or online book.

7.  Open bibliographies and reviews.  Let your colleagues know what you're reading, and what you've learned from it.

Thursday, November 10, 2011

Book Review: Reinventing Discovery

I believe that the process of science—how discoveries are made—will change more in the next twenty years than it has in the past 300 years. --Michael Nielsen, Reinventing Discovery
I appreciate an author who's not afraid to make bold claims, and Michael Nielsen certainly fits that description.  He goes on to say even that
To historians looking back a hundred years from now, there will be two eras of science: pre-network science, and networked science.  We are living in the time of transition to the second era of science.
I grew up feeling that the golden age of science was the first half of the twentieth century, which gave us marvelous advances like relativity and quantum mechanics.  According to Nielsen, though, I'm witnessing the most transformative period of scientific development since the invention of the scholarly journal in the 1700's.  Although I'm a firm believer in the power of the internet to accelerate scientific advances, I was skeptical.
I downloaded Michael Nielsen's Reinventing Discovery on Tuesday and read it in less than 48 hours (between shopping trips while on vacation in Dubai).  Although I was familiar with much of the material in the book, it was an engaging and highly thought-provoking read that I think both scientists and laypersons will enjoy.  I'll focus here on the ideas that struck me as especially insightful.
Nielsen gives several examples to illustrate the beginnings of his foretold revolution; some are scientific (the Polymath projectGalaxyZoo, FoldIt) while others simply illustrate the power of our new networked world (Kasparov versus the World, Innocentive).  These examples are used extensively and lend a convincing empricism to a book that claims to predict the future.  They also allow Nielsen to dive into actual science, adding to the fun.
Many scientific advances are the result of combinations of knowledge from different fields, communities, or traditions that are brought together by fortuitous encounters among different people.  In a well-networked world, these encounters can be made to happen by giving individuals enough accessible information and communication. Nielsen refers to this as "designed serendipity".
The reason designed serendipity is important is because in creative work, most of us...spend much of our time blocked by problems that would be routine, if only we could find the right expert to help us. As recently as 20 years ago, finding that right expert was likely to be difficult. But, as examples such as InnoCentive and Kasparov versus the World show, we can now design systems that make it routine.
Offline, it can take months to track down a new collaborator with expertise that complements your own in just the right way. But that changes when you can ask a question in an online forum and get a response ten minutes later from one of the world’s leading experts on the topic you asked about.
The trouble is, of course, that the forum in question doesn't exist -- and if it did, who would have time to read all the messages?  Nielsen delves into this question, discussing how to design an "architecture of attention" that allows individuals to focus on the bits most relevant to them, so that large groups of people can work on a single problem in a way that allows each of them to exercise his particular expertise.  Taking the idea of designed serendipity to its logical yet astounding conclusion, Nielsen presents a science fiction (pun intended) portrayal of a future network that connects all researchers across disciplines to the collaborations they are most aptly suited for.  I found this imaginary future world both fascinating and believable.
The second part of the book explores the powers that are being unleashed as torrents of data are made accessible and analyzable.  Here Nielsen draws examples from Medline, Google Flu Trends, and GalaxyZoo.  While the importance of "data science" is already widely recognized, Nielsen expresses it nicely:
Confronted by such a wealth of data, in many ways we are not so much knowledge-limited as we are question-limited...the questions you can answer are actually an emergent property of complex systems of knowledge: the number of questions you can answer grows much faster than your knowledge.
In my opinion, he gets a bit carried away, suggesting that huge, complex models generated by analyzing mountains of data "might...contain more truth than our conventional theories" and arguing that "in the history of science the distinction between models and explanations is blurred to the point of non-existence", using Planck's study of thermal radiation as an example.  Planck's "model" was trying to explain a tiny amount of data and came up with terse mathematical equations to do so.  The suggestion that such a model is similar to linguistic models based on fitting terabytes (or more!) of data, and that the latter hold some kind of "truth" surprised me -- I suspect rather that models informed by so much data are accurate because they never need to do more than interpolate between nearby known values.  Nevertheless, it was interesting to see Nielsen's different and audacious perspective well-defended.
A question of more practical importance is how to get all those terabytes of data out in the open, and Nielsen brings an interesting point of view to this discussion as well, comparing the current situation to that of the pre-journal scientific era, when figures like Galileo and Newton communicated their discoveries by anagrams, in order to ensure the discoverer could claim credit later but also that his competitors couldn't read the discovery until then.  The solution then was imposed top-down: wealthy patrons demanded that the discoveries they funded be published openly, which meant that one had to publish in order to get and maintain a job.
The logical conclusion is that policies (from governments and granting agencies) should now be used to urge researchers to release their data and code publicly.  Employment decisions should give preference to researchers who follow this approach.  At present, the current of incentives rather discourages such "open science", but like Nielsen I am hopeful that the tide will soon turn.  I was left pondering what I could do to help; Nielsen provides numerous suggestions.  I'll conclude with some of the most relevant for computational scientists like myself.
...a lot of scientific knowledge is far better expressed as code than in the form of a scientific paper. But today, that knowledge often either remains hidden, or else is shoehorned into papers, because there’s no incentive to do otherwise. But if we got a citation-measurement-reward cycle going for code, then writing and sharing code would start to help rather than hurt scientists’ careers. This would have many positive consequences, but it would have one particularly crucial consequence: it would give scientists a strong motivation to create new tools for doing science.
...
Work in cahoots with your scientist programmer friends to establish shared norms for citation, and for sharing of code. And then work together to gradually ratchet up the pressure on other scientists to follow those norms. Don’t just promote your own work, but also insist more broadly on the value of code as a scientific contribution in its own right, every bit as valuable as more traditional forms.

Thursday, November 3, 2011

Collaborative scientific reading

I often feel that the deluge of mathematical publications, fueled by the ever-increasing number of researchers and mounting pressure to publish, threatens to overwhelm my ability to keep up with advances.  I don't think this is peculiar to applied mathematics.  No matter how adept you are at sifting the chaff and finding the most relevant work in your field, you won't possibly have time to read every paper that is germane to your research, let alone those of tangential interest that might provide new research avenues.  For my part, although I take time to read new papers every week, I've resigned myself to the fact that I won't see more than the abstract of most of the papers I'd like to read, because I need to conduct new research, teach, write, and so forth.

Reading and digesting a mathematical paper takes time and concentration.  Nevertheless, I find that perhaps 80% of the value I get out of reading most papers can be summed up in a paragraph or two that is easy to read and understand.  We all have practice producing those terse paragraphs because we regularly referee papers and provide a concise summary for the editor.  This summary includes things like "what's really new in this work" or "how this relates to previous work", as well as an evaluation of its merit.  Unfortunately, those referee reports are kept secret and unavailable to our colleagues.  I mentally create a similar report for most papers that I read in depth, although I don't usually write my evaluation down and I certainly don't send it to anyone.  What if every reader of a paper had access to the summaries and evaluations made by all the other readers?  I think we could all learn a lot more, a lot faster, about what our colleagues are accomplishing.

Recently, Fields medalist Timothy Gowers proposed an approach to accomplishing just that. The idea is to bring the functionality of StackOverflow to the arXiv, creating a place where everyone can publish and everyone can openly referee or comment.  The StackOverflow system of reputation and up-/down-voting would be used to help the best papers and best comments float to the top.  As Gowers admits, there are plenty of obstacles, but I'm hopeful that people with his level of clout in the mathematical community could really bring this to pass.  His interest seems mostly based on issues with the current journal publication system, but I see it primarily as a way to "collaboratively read" the literature.  Indeed, it might be best if the site had no implications for decisions on hiring or tenure, to avoid any motivation to game the system.  The site would also be a great place for expository writing that can't be published in a journal.

It's encouraging to see that some things are already moving in this direction.  A new website named PaperCritic has just been launched to accomplish something roughly along these lines.  It doesn't involve the StackOverflow system, but has Mendeley integration and allows you to post a public review of any paper.  Meanwhile, an increasing number of scientists are including paper reviews in their blog posts -- something I would like to do here.

I think Mendeley could accomplish something useful in this direction if they would give users the option to make their library and notes public.  Then when I find a paper on Mendeley that says "20 Readers", I could find out who they are, see what they've written about that paper, and see what else they're reading.

Note: I know that we already have Mathematical Reviews, but in my opinion it doesn't accomplish the goals mentioned above, mainly because the reviewer of a paper is often not sufficiently knowledgeable about the paper to say anything more insightful than what's in the abstract.  I find that Mathematical Reviews gives me papers to review that I would never have read otherwise.  What I'd like to see are reviews from the people who read the paper because it's germane to their own work.

I discovered while writing this post that there was until very recently a successful site of this kind used by quantum computing researchers called scirate.com.  Perhaps we should focus on helping this guy get the site back up and start using it for math too.

Edit: Another brand-new open review system: http://open-review.org/

 

Thursday, May 19, 2011

What is science?

Today I received an e-mail from a collaborator of mine stating

...our community doesn't reward engineering effort but instead scientific advances.

The context was a discussion of what is publishable in scientific software development. This got me thinking about what really is the difference between science and engineering, and what is the difference between publishable advances and the rest of the work that scientists do.

After some thought I've concluded that this difference is in many respects artificial and purely a function of one's discipline (or even sub-discipline). To take algorithmic complexity as an example:

-To a complexity theorist, a reduction in complexity from 1000N^3 to 3N^2 would be considered "engineering effort": after all, either one is in P, right?

-To an applied mathematician, the above result would be considered a fabulous scientific advance, but a further reduction from 3N^2 to 2N^2 might not be deemed publishable.

-To a chemist, say, who needs to run the algorithm with 10000 different parameter sets, even a 10% improvement might be considered a valuable scientific advance.

Of course, to a pure mathematician any algorithm for solving the problem is an engineering detail; the only scientific aspect is proving the existence of a unique solution.

All of these advances may be important steps leading from an abstract idea to a technological breakthrough that benefits non-scientists. I think it's okay that all the above researchers are only interested in one particular step in this chain of advances. What's detrimental, however, is the inability to recognize that the other steps in this chain are valuable. This kind of bigotry is, in my experience, rather common among scientists.

Edit: Ironically, the quote at the beginning came from a person whose title is "Professor of Engineering".