Thursday, 14 January 2010

Men at Work

As soon as I began experimenting with the DOSY toolbox I was surprised to discover an unexpected hole, which is just a novel proof of my old theory on NMR imprinting.
If your first spectrometer is a Varian, you will become a certain kind of spectroscopist and a certain kind of programmer. If your first instrument is a Bruker, you will become a different spectroscopist and a different programmer.
If a brukerist programmer had written this toolbox, he would have imported the frequency domain spectrum and started from there. Mathias Nilsson is instead a Varianist and he imports the FID only. We talked about it and now he agrees with me on the importance of importing and exporting the data to and from as many programs as possible. For example, I might prefer using an external software for the preliminary processing (FT, phase and baseline correction), then the DOSY toolbox for the separation of components, eventually a third software (or the first one again) for printing, visualizing, creating slides, etc...
This is not really efficient, because in many real cases I probably want to try with different preliminary processing (e.g.: apply or remove a line broadening), and this is already easily done with the toolbox, but becomes cumbersome if I move data from an application to another at each stage. The principle is, however, to let the user choose whatever he prefers.
This kind of transfer is the less problematic one: the information is not to be stored, not to be transmitted, is consumed on the spot on the same machine that generates it. I and Mathias have already started collaborating on this and, if only we were better programmers, we had already finished by now, but this is not the case.
Unfortunately, we are equipped with two different brains, and it has a price: countless emails to explain each other what we need, what we can do, how we would solve a given problem. It's really a case where it takes more to explain a thing than to do it.
In practice, we have created a simple, highly readable, redundant file format that consists into a text file, like JCAMP-DX, but without compressions of any kind. At the same time, we are testing this format on the field. I am writing two scripts (the language is Lua) to export data from iNMR to the DOSY toolbox and to import data back again.
We have chosen this approach because I am the prototype of the average ignorant. If I am able to write these scripts for at least one external program, this can be categorized as an easy task, whichever other program you want to use in combination with the DOSY toolbox, the only requisite being that the program must be scriptable.
Now you know what I have been doing during these days and why the promised review is not ready yet. To tell the truth, I don't know why I should write a review in a case like this, where the source is open and the author is so collaborative. I mean: if I find something that I don't like, it's more productive to talk about it directly with Mathias Nillson than to spend a day to write a review.
Generally speaking, I don't believe in open source, I don't see how it might work. Pick any example of a moderately large project: how many people can safely modify it? Normally, it's only the original author. In practice, it makes little difference if the source is open or not.
In this particular case, the software is written with MATLAB, which I haven't, but it would be the same if it was written in a language I am familiar with.

To be continued...

Thursday, 7 January 2010

Installing the DOSY toolbox on a Mac

The DOSY toolbox is an open source program written by Dr. Mathias Nilsson that runs on Windows, Linux, Mac and on any other platform covered by MATLAB. I am going to review it next week, while today I am giving a few practical tips that I prefer to leave out of the review, for the sake of readability and tidiness.
I know that 14% of my readers have a Mac, and I have verified that the original installation instructions of the toolbox are inaccurate and unclear (for the average Mac user at least), so today I will simply explain how you can install the DOSY toolbox on your Mac.
Download the file called DOSYToolbox06_MAC_13May09_pkg from the address:
http://personalpages.manchester.ac.uk/staff/mathias.nilsson/software.htm
With a double click this archive generates a folder called: "DOSYToolbox06_MAC_13May09_pkg 2". This name is actually too long and useless, therefore I have shortened it into "DOSYToolbox06" and moved the folder inside my home directory. This is not strictly necessary but is a safe and sensible move. Inside the folder there is an installer called "MCRInstaller.dmg". It installs the runtime of MATLAB, which is necessary (unless you already own MATLAB itself). The difference is... technically speaking... to make it short... well the difference is $2,000 that you save if you install the runtime instead of buying matlab itself. Got it? A little slower but much cheaper.
The installer works like any other installer: you must always say "yes", you can't be wrong, eventually you will read it's all OK but you have no clue of what it has done and this is terribly bad because you absolutely need to know where the files have been installed, otherwise you can't run the DOSY toolbox at all. Fortunately I have discovered where the files have gone:
/Applications/MATLAB/MATLAB_Compiler_Runtime/v710
(this is valid, of course, for version 7.1 of the runtime).
The downloaded folder also contains a script called: "run_DOSYToolbox06_MAC.sh". This is the command that starts the program, and it needs an argument and this argument is the path to the matlab runtime that you have already installed. In other words, the canonical way to start the toolbox would be to type a command like this into the Terminal:
/path/to/run_DOSYToolbox06_MAC.sh /path/to/MATLAB_Compiler_Runtime/v710
for example:
/Users/your_name/DOSYToolbox06/run_DOSYToolbox06_MAC.sh /Applications/MATLAB/MATLAB_Compiler_Runtime/v710
This is highly impractical, but you can easily create a double-clickable shortcut:
- open the Automator application
- choose the template called "workflow"
- into the leftmost column, click on "Utilities"
- into the second column, double click on "Run Shell Script"
- from the pop-up menu "Shell" choose "/bin/bash"
- fill the main field with our command:
Run this workflow. If, in a few seconds, the DOSY toolbox window appears, it means that the workflow works. Quit the DOSY toolbox and save the workflow. Be careful and select the proper File Format: it should be "Application" and not "Workflow".
Finally you have created a double-clickable application (you can paste a different icon)
and you can forget Terminal, shell, bash and the likes...
At the time of this writing, though, it seems that the DOSY toolbox is fully functional only if launched from a shell. It is a puzzling and unwelcome incident, I hope it's going to find a solution soon.
Has anybody the compiler for the Mac? Would this person be so kind to compile version 0.7 for the rest of us Mac users? "Open source" means that some times you should also give, not always take.

Wednesday, 2 December 2009

Jake Bundy

What's your position and where are you working?
For the past 5 years I have worked as a lecturer in Biological Chemistry in the section of Biomolecular Medicine at Imperial College London.

Where have you been working before?

Before this, I post-doc’ed at Cambridge; before that, UC Davis; and before that, for my first post-doc, at Imperial College again.

Briefly describe your research.
I am interested in metabolism in invertebrate and microbial species, and how this is involved in several different biological questions. Some of the projects I currently work on include microbial virulence and pathogenesis; how metabolism is affected by problems with recombinant protein folding in the bioprocessing yeast Pichia pastoris; and using earthworms as biomonitors of environmental pollution.

What do you use NMR for?
Together with mass spectrometry, I use NMR for metabolite profiling, as part of the technology for metabolomic studies. Although it’s generally less sensitive than mass spec, NMR still has a very useful role – as a near-universal and robust detector, it can give a quick and information-rich spectral profile. Most of the work we do, we just use 1D NMR for profiling; particularly useful for studies where you want to process as many samples as possible. However there are also other cases where more in-depth NMR experiments are needed, for example for isotopomer analysis to investigate metabolic fluxes; or to assign novel metabolites (essentially a natural products chemistry problem).

Which NMR software are you using?
XWIN-NMR and Topspin for data acquisition; iNMR for all NMR processing. I also use Chenomx NMR Suite for helping assign and quantitate metabolites in NMR spectra.

Which other NMR software have you used in the past?

I’ve also used VNMR and ACDLabs NMR software, and MestreC (before it was released as a commercial product).

How do you rate iNMR?
iNMR is not only my favourite NMR software for Mac OS (I’m not sure how many competitors there are at the moment), but it’s easily my favourite NMR software full stop. It’s one of a handful of Mac-only packages that I use all the time as part of my regular working day (others include Papers, Aabel, and Bookends). Features that I particularly like include the Overlay Manager (which makes it by far the quickest and easiest NMR software to use for comparing multiple spectra, in my opinion), and also the overall simplicity of using it to produce quality spectral images that can go straight into a paper or presentation without having to use multiple machines or virtualization. I admit I wasn’t really bothered by lack of anti-aliasing on spectra before using iNMR, but now I’m used to it, I do find it genuinely annoys me when looking at spectra in Topspin, say – it’s distinctly harder to see fine detail without zooming in. It’s also elegant and quick – not essential properties for software, but makes it more enjoyable to use on a regular basis.

Is it enough for your needs?
Well, it’s certainly enough for my needs in the sense that all of the spectra that I acquire are processed with iNMR – so in one sense, yes, almost by definition. It’s not 100% perfect though, there are still some small issues that could be ironed out in future releases – and as I’ve already said, I do use Chenomx software for some complementary uses which iNMR isn’t primarily designed for. I definitely see it as a crucial part of my workflow for the foreseeable future though, and expect it will keep improving (although by now it’s a relatively mature product).

Hands on 3-D Processing

A measure of the computing power available today with a desktop computer is the possibility of processing huge spectra in real time. For example: is it possible to correct the phase of a 3-D matrix interactively, in real time? The answer is: yes and you don't even have to employ more than a single core nor to buy an high-end graphic card. When I mean interactive processing I mean that:
- you see a graphic representation of the matrix at each processing stage.
- you can play with it, for example change a parameter just to see the effect it has on the matrix.
It is always necessary to know and understand the mathematical rules that govern NMR processing. The more you know them the more you enjoy their visual representations. The more you play with the graphics, the more you understand the maths behind. The two things go together.
You can find a self-teaching course on basic 3-D processing on the iNMR web site. There is no theory, only a lot of examples and a good measure of practical tips.

Thursday, 26 November 2009

Open Source NMR freeware

Most of the readers arrive here using Google, without knowing me and my blog. Usually they get very angry because they arrive... on the trapping post I wrote 3 years ago! I want to do something to keep them glad...
So you want "open source" stuff? Do you know what it really means? Are you ready to compile, test, debug it and add a graphic interface to it?
Just because you asked for it, here is a list of available projects. If you know other links, add them into a comment.

CCPN
NPK
matNMR
ProSpectND
Connjur
Newton-NMR
nmrproc
DOSY Toolbox
list of 30+ projects

Monday, 16 November 2009

Tips

It rarely happens to find valuable tips about processing on the web. When it happens, it's probably not enough to bookmark the page (it may disappear), copying it is a better idea. Here is the link:
http://spin.niddk.nih.gov/NMRPipe/embo/
When you arrive there, scroll down until you find the chapter Some General Tips About Spectral Processing.
The focus is on multi-dimensional processing with NMRPipe, but a few concepts are generally applicable indeed.
The brief discussion about first-point pre-multiplication is something to bear in mind. You can also find clearly expressed opinions on zero-filling, linear prediction and baseline correction.

Saturday, 14 November 2009

Shadow


Here is a picture I have found on the internet. I have never been in this place, if this is what you want to hear. Suppose, instead, that I live just in front of this tower and tomorrow I take a photo of it immediately after dawn, then another photo after 25 hours and so on for a week, with regular intervals of 25h between pictures. Eventually I print all the pictures in order and ask you:
"This pictures have been taken in this exact order at regular intervals. Can you tell me how long the intervals were?".
Somebody will answer: "1 hour" and the answer would be partially correct. A more correct answer would be 1 + 24 n hours, with n = 0, 1, 2, 3….
The world of FT-NMR is similar. The difference is that the regular interval between consecutive observation is known in advance while the speed of the hands and of the shadow is unknown. In other words, it is the opposite of my example of the tower.
The equivalent of a day, in the world of FT-NMR, is very very short and is called dwell time. It forms a Fourier pair, so to speak, with the spectral width. The spectroscopist sets the former, the latter is a mere mathematical consequence.
The ignorant says that the spectral width is the distance from the first point of the spectrum to the last one. This statement is as incorrect as saying that a day is made of 23 hours! The spectral width, actually, is the distance from the first point to the first point after the last one!
Let's verify it with a numerical example. Let's say we have a spectrum of 1024 points, separated by 2 Hz. If we zero-fill the FID up to 2048 points, the distance should decreases to 1 Hz. The spectral width is 2048 Hz in both cases, and this is OK. If you measure it the other way, then you have a spectral width of 2046 Hz that grows up to 2047 Hz. This is absurd, because the value is fixed at acquisition time and can't be changed by processing.
The relation is Spectral_width = 1 / Dwell_time.
Many books report another formula: Spectral_width = 1 / (2 * Dwell_time). This assumes an instrument without quadrature detection, in other words with a single detector. I have never used such an instrument.
To be exact, my whole description is dated, because today's instruments work in oversampling, the actual dwell times are shorter than what the spectroscopist sets (4 or 8 times the value reported), the FID that we see is already the result of a couple of FT (first direct, then inverse), et cetera. Seems complicated but it is not. We see what we need to see, the complications are hidden.
The important concept to remember is that the components of a digital spectrum, even if they are called "points", should be treated and conceptualized as tiles. There is no room in between. This idea will help you when you'll try to save a spectrum as a table of intensity vs. frequency. You will get the correct frequency value of each "point".