Wednesday, 18 August 2010

Can Zero-Filling Correct the Baseline?

I want to show you a proton spectrum that has puzzled me during the last weeks. It contains something that's quite typical and something that I can't explain. I have processed the spectrum in two different ways, with zero-filling and without it. The spectrum without zero-filling is black, the spectrum with zero-filling is green (the number of points is doubled).
This detail is the bottom part of the TMS signal (magnified to show the ringing effect). Where does the ringing come from? TMS is a small symmetric molecule and its protons have a long relaxation time. Their signal persists at the end of the FID. When we add the zeroes after the signal, a step is created. The FT of the step is the ringing that we see. The spectrum without zero-filling doesn't contain the step, so there is no ringing. The period of the ringing is exactly 1 point. In simpler words: odd points are positive, even points are negative. Things are not so simple, actually, because the rule is reversed on the two sides of the peak. This is something I have always seen, I don't know if it's a constant rule or something that's merely more probable than its opposite.
Without zero-filling, we have only half of the points. They correspond to the maxima on the left of the peak and to the minima on the right of it. Any program for automatic phase correction is fooled by asymmetric peaks like this. Even humans are often fooled. They think that the spectrum is "difficult to phase" and don't recognize that the peak is asymmetric. Asymmetry and ringing are two sides of the same coin. Without zero-filling we have asymmetry, with zero-filling we have ringing. In the first case it is difficult to recognize that the signal is truncated, because it appears much larger than it actually is.
Up to this point I can explain everything. It's all familiar to me. There is another effect that I can't explain at all and appears when I observe the whole spectral range. The baseline of the normal spectrum is wavy.
The baseline is perfectly flat in the other case. This is the first time I see such an effect: can zero-filling correct the baseline?
This spectrum was acquired on a recent Jeol 400 MHz instrument. I wonder if the digital filter has anything to do with the latter effect.

Friday, 6 August 2010

Which type of barista are you?

A colleague from New Zealand once was telling me about a type of coffee that was originated there. It's called Flat White Coffee. We were discussing about the difference between this coffee and all the other types of coffee with milk. Eventually we started talking about the quality of the drink and what makes it be better or worse. He mentioned that, of course, the quality of the beans and milk are very important for a good coffee but what makes a good coffee an excellent coffee is the barista's ability.

What makes a coffee to be a rubbish coffee then? If the coffee grains are rubbish, you will have a rubbish coffee. However, for a flat white coffee, coffee is just one variable in the equation. There are other variables like how the grains are roasted, steaming the milk at the right temperature, not adding sugar, how the milk is poured, the microfoam on top of the drink, etc. See the distinction from cafe con leche for details. Anyway, the point is, a rubbish (or careless) barista, with rubbish coffee will produce a rubbish coffee drink.

Rubbish Coffee
Clearly a business that serves a rubbish coffee will not survive if their main business is to sell coffee drinks. Also, this business will never attract customers that really appreciate a good coffee.

Average Coffee
Producing an average coffee is easy. You buy average coffee grains, hire an average barista and, hey presto, you have an average coffee drink. There is nothing wrong with an average coffee drink if you are not in the coffee business. However, if you are in a coffee business, your business will be just another one. No one will remember you and chances are that you will have occasional customers, but not regulars.


Good Coffee
Producing a good coffee is not that simple. However, it does not need to be expensive. You don't need to buy the best quality ingredients to be able to make a good coffee. Best quality ingredients are expensive and inevitably will make your coffee drinks more expensive as well. You can mitigate this situation hiring a good barista, or maybe someone that has the potential and willingness to become a good one. Clearly, the barista that prepared the coffee above is trying his best to make a good coffee. It's still not perfect (purely looking at the pattern on top of the drink) but it is definitely a much better coffee than the normal and average one that you get everywhere. It is clear that the barista cares about it and eventually he will be able to produce a very good coffee.If you are in the coffee business, producing anything less than a good coffee is just unacceptable.

Great Coffee
And then you have the great coffee, that is made with great coffee beans, carefully roasted, prepared in a very good coffee machine and by a great barista. Great baristas are proud of their ability to prepare a great coffee and they would not work for too long for a business where coffee making is not treated with the deserved respect.

Now imagine that you are the barista. But instead of coffee, you produce code. Imagine that the the coffee grains, milk and coffee machine are the tools you use to produce your code like the computer language, the IDE, the database, etc. Instead of serving coffee for your customers, you are producing software that will be used by your team mates, project sponsors, the company they (or you) work for, external clients, etc. 

Like in the flat white coffee example above, of course that the tools we use are very important when producing a good software. However, more importantly, it is the quality of the software engineers that counts. A good barista can make a good coffee even when using average coffee grains, due to his or her ability to combine ingredients and prepare the drink. A bad barista can ruin the coffee even if he or she is using the best quality coffee beans. The biggest difference between the two baristas is how much they care about each cup of flat white coffee they prepare. Their pride and willingness to achieve the best pattern on top of each drink. The great feeling of achievement when they produce a great one. The happiness to see returning customers, queuing and waiting their turn to order the coffee that he, skilfully, prepares.

So, which type of barista are you?

Saturday, 17 July 2010

The Barrier

In the last two years NMR software has kept evolving slowly, but the world around it is changing more rapidly. From a technical point of view nothing important happened; while from a commercial point of view we are in the middle of a revolution. Prices have dropped down considerably. The false categories, the so-called "professional", "industry-standard" programs on a side and "low-budget", "alternative" programs on the other side, have disappeared. The companies that were selling the products in the first category have realized that the true values were actually reversed and adjusted the prices accordingly.
Consider TopSpin, for example. I remember its original price was in the range € 4000-5000. At the beginning of the month, I have visited an industrial lab where they had recently purchased a license for off-line processing. They told me they had paid € 3000 for TopSpin. They had also considered a solution by ACDlabs, but after receiving a quote of € 10,000 they renounced with disgust. The main reason why they prefer TopSpin is that it is possible to process the spectra directly on the spectrometer, with great saving of time. When back into the lab, they simply print the already processed spectra.
Less than 24 hours later I had a pleasant talk with Angelo Ripamonti, from Bruker Italy. He gave me different figures. The price of a license is € 2,000, complete with box, cartaceous manual and CD. They also have an electronic edition, called "Topspin Student Edition". It is the same product, the license expires after 3 years and costs € 99 (as far as I understand, there is no after-sale support; I am not sure, though). With my great surprise, Angelo said that they hope the students will familiarize with TopSpin during their PhD period and remain faithful to it. I used to think that TopSpin was by far the most famous of all the NMR programs; evidently they are not so sure and fear the competition.
On one thing we agreed all along the line: the chemical industry has completely disappeared from Italy (if we are allowed to measure it by the number of magnets being sold). If I could, I would also close the Italian universities (what's its purpose, if there is no industry? the local job market wants nurses, not chemists), but that's another story (Angelo hopes the Universities will grow and buy more magnets).
I told Angelo that I had noticed that ACD had their own "academic edition", which is completely free. I myself have received a copy, though never used yet. I asked for his opinion: could the ACD move be a reply to Bruker's Student Edition? "No", he said, "they are making a lot of money from their NMR database". Translation: ACD is the only competitor in the field of NMR database and they want to monetize as much as they can while the favorable situation persists. Indeed, the "price" I paid for the NMR processor was my email address, which they have already used to advertise the database.
Bruker is also lured by this segment of the market. They bought Perch and are working on it to make a new product that will predict the chemical shifts from the structure. Not the same thing as a database, but with the same purpose.
As my readers know, I had long waited to try the ACD program, because I started this blog with the intention of writing reviews for all the software in existence. Time passed and I have different pastimes today. If I haven't found the time yet to study the Processor and write a review, there are very few chances I can do it in future. The mere length of the manual discourages me. I want however to comment on their commercial move.
(1) It is not correct to first sell a program to several universities and then to give it for free to the rest of the world. Unless they give the database for free to those universities that paid for the processor. That would be a fair compensation. What about, however, the universities who already have bought both products? What about those who never bought anything? After this precedent, how can ACD convince somebody to pay for one of their products, when there's the risk that it becomes free after a few years?
(2) People see with different eyes a program that is free starting from version 1 and a program that is free starting from version 12. Consider, for example, Internet Explorer. It was born as a freeware and it was a huge success, actually a monopoly. It still accounts for 45% of the market. Firefox accounts for the 32%. Opera, instead, started as a commercial product before becoming free. Its share of the market is a disappointing 2%. What people feel is that if the former commercial product was not good enough to sell copies, it is not good enough to bother with.
(3) I don't know exactly why, but ACD has not been lucky, so far, with free programs. Consider their ChemSketch, for example. I have never found a bad review, while I always hear people complaining about ChemDraw. Somehow the latter remains the undisputed no. 1 in the field. Why?
(4) If you want your product to become popular, making it free is a good move, but giving it a catchy name is as much as important. People at ACD have always lacked in inspiration. "ACD NMR Processor Academic Edition" is the less memorable of the names. Even a distasteful name like "BottomSpin" would have proved more effective.
(5) If you analyze the situation in detail, this product is still far from being "free". Whenever you print, the program adds a red line reminding it's not to be used for commercial purposes. The file format is proprietary and not recognized by all the programs (you are locked-in). There is a potential risk that the program becomes commercial again in future.
The battle has reached the first result. It has created a barrier. If a new competitor wants to enter into the arena, its product must be at least as good as today's free programs (and there are at least 3 formerly commercial programs that are available for free today). Given that it costs a lot of time and money to create a new NMR program, while the prices keep falling, the barrier is very high. We are not going to see any novel program in the next decade, neither free nor commercial. If all the investments are made in the field of databases and predictors, even existing processing programs are not going to be revamped, but simply refreshed.
Do not think that users are happy and programmers are sad. Not at all! After 4 years, the most read page of this blog remains "TopSpin Free Download" and readers who comments there are angry as before (the page exists, but there is nothing to download!). All the programmers I am talking with, on the other hand, are quite glad: their programs are selling and nobody went in bankrupt in the last few years, despite the financial crises. Apparently a point of contact has been found between the two sides. The prices are more reasonable than they used to be a decade ago, the programs are better and many customers prefer to pay if they can receive a good service in addition to the product.

Thursday, 1 July 2010

Universal Hole

Earlier in this week I cited a paper by Kobzar and Luy. It contains the statement:
The coupling extraction procedure is not yet implemented in any available software.

that confirms and enforces what I have always being saying:
Any NMR program contains some hole and by the time it's filled another hole appears.

What they have found is a kind of super-hole that is common to every program. My first thought would normally be: "If nobody cares, why should I?", but this time I was intrigued by a figure just above the cited statement. That figure resembles a picture of mine I published here a few months ago. Despite the apparent similarity, however, the two methods have little in common.
Driven by curiosity, I looked on the web for anything more recent on the same subject and found this page that describes the very same "long range J" procedure. Does it mean that somebody has already filled the hole?
I have contacted the PR man at nucleomatica and he explained that the procedure is not commercial yet. It is a very simple data manipulation, there is no secret about it, but neither there is demand for it by the market. In conclusion, there is no hurry to make it available (to a distracted public).
Finally he gave me this picture, which is a world-exclusive of my blog:
Believe it or not, what you see is the same multiplet shown into the JMR figure (page 133, fig. 3d). Same molecule, same kind of experiment, another sample, another instrument.
The two experimental multiplets in black differ for the absence (top) or presence (bottom) of an anti-phase coupling. Both traces come from 2-D experiments and the resolution can never be enough to directly measure the size of the coupling.
The green circle hilights a slider. When you move the slider, the program adds an artificial coupling to the upper trace. The result is shown in red. When the red multiplet is like the black multiplet at the bottom we have succesfully simulated the missing coupling AND NOW WE KNOW HOW LARGE IT IS. That's what it's all about.
If you remember, I have gone much further with my unbeatable simulator, because it is able to extract all the couplings with a single experiments.
The strenght of my method is, alas, also its drawback: even when you are interested into a single J value, you are forced to measure them all. It can be very hard in cases like this.
Despite the external similarities the two methods are very different inside, serve two different purposes and can live side by side very well into the same program.

Tuesday, 29 June 2010

Faster Faster Faster

Our machines, even when hitting an apparent performance peak, only run at one small fraction of their true potential speed. I feel that today's computers and their software are OK for routine spectra. I couldn't ask for more. Other spectra are quite large, however, and I must wait a few seconds during processing. Without going into the third dimension, consider these novel experiments to measure long range heteronuclear Js. Each row contains at least 4096 points. Quite likely we are going to see larger rows in the next few years. The time required to compute the FFT is in the order of the seconds. It would be great if we could half this time. The solution is public since 2008 and freely available. It must also be well-know, because I have initially found it on Wikipedia.
What they say, in practice, is that if you use the GPU (the graphic chip) instead of the CPU (the main brain of the computer)...
In this work we present a novel implementation of FFT on GeForce 8800GTX that achieves 144 Gflop/s that is nearly 3x faster than best rate achieved in the current vendor’s numerical libraries.

Another paper says:
We implemented our algorithms using the NVIDIA CUDA API and compared their performance with NVIDIA's CUFFT library and an optimized CPU-implementation (Intel's MKL) on a high-end quad-core CPU. On an NVIDIA GPU, we obtained performance of up to 300 GFlops, with typical performance improvements of 2--4x over CUFFT and 8--40x improvement over MKL for large sizes.

The source code (to be compiled), is available on another site. What upsets me is the ReadMe file:
Currently there are a few known performance issues (bug) that this sample has discovered in rumtime and code generation that are being actively fixed. Hence, for sizes >= 1024, performance is much below the expected peak for any particular size. However, we have internally verified that once these bugs are fixed, performance should be on par with expected peak. Note that these are bugs in OpenCL runtime/compiler and not in this sample.

Maybe they have already found the solution to this problem.
There's another issue, however: how long does it take to move the matrix from the main memory to the GPU and back? I presume that loading the columns will take much more time than loading the rows. In this case it is better to transpose the matrix between the two FFTs (just like when we use the CPU). Eventually, the bottleneck will the the transposition, not the FT.

Wednesday, 23 June 2010

One team, one language

On a previous post I was discussing, among other things, how code often doesn't represent the business properly. The code "satisfies" the business requirements but doesn't express them very well. The main reason for that is because we, developers, like to abstract business terms and rules into technical implementations and patterns.

Very often, we discuss the user stories (requirement documents, use cases, whatever the methodology used is) with the "domain experts" and as soon as we understand what needs to be done, we map the requirements to a technical design (actions, services, entities, helpers, DAOs, etc) that is completely meaningless to the domain experts. Sometimes, even among developers themselves, different names and expressions are used to refer to the same thing. The main reason is that different developers talk to different domain experts and come up with different abstractions. As a result,  the usage of different terms to describe requirements leads to confusion, duplication of code and unpredictable behaviour in the system.

The first step towards an expressive and domain-focused design is to have a common language among ALL members of the team. ALL means ALL: developers, domain experts (business analysts, users, product owner, etc), testers, project manager and anyone else involved in the project.

Developers very often say that domain experts don't understand objects and database and because of that, they need to "translate" business requirements into software design. However, we developers don't understand the business as well as the domain experts do, what more often than not, leads to imperfect and confusing abstractions.

The Ubiquitous Language

A language structured around the domain model and used by all team members to connect all the activities of the team with the software.
The Ubiquitous Language is one of the most important things, if not the most, in Domain-Driven Design. The main idea is that the whole team speaks a single language, that is the business language.



Business terms related to the software to be implemented must enter the ubiquitous language and each of these terms must be understood clearly by all members. During requirements gathering sessions and planning meetings, these terms must be captured and made available to everybody. Technical terms from the development team and business terms not relevant for the piece of software being implemented MUST NOT enter the ubiquitous language. 

Capturing the ubiquitous language

In order to capture the ubiquitous language, it is mandatory that you work on a iterative software development environment. Trying to capture the ubiquitous language up-front could straitjacket the whole process, inhibiting team members to make the necessary changes along the way. As the language is used to express the business requirements, it is natural that it evolves during the lifetime of the project, where new terms are added, deleted and also re-defined.


Methodologies like Extreme Programming (XP) says that the only documentation should be the code. Not even comments on the code are appreciated, since they can easily get out of sync with the code. The code should be the only documentation since it is the only one that represents exactly what the system does.

On the other hands, we have UML (Unified Modeling Language), that in theory, should be a great candidate to document the ubiquitous language since the whole purpose of UML was to document requirements and express them in a language that is common to developers and business people.

In summary, showing code during discussions with domain experts, testers and other members of the team during a design session is not exactly a fantastic idea. Also, using just UML, because of its bureaucracy, rules and details is also a bad idea. The whole UML notation could easily straitjacket the creative process during the exercise.

There are many discussions about what would be the best way to capture the ubiquitous language. My preferred way is to draw diagrams (boxes and arrows mixed with some well understood UMLish notation) where each box represent a "domain object" (aka domain concept). A domain object can be anything that is expressed by the business, like client, organisation, product, route specification,  sales system, etc. They would all be boxes. Add to it a few arrows linking the boxes and with just a couple of words explaining how they related to each other. Sometimes a mixture of a class and sequence diagram (or an active diagram) can be very helpful, but don't get to picky about any notation. Preferably, draw on a white board, take a picture and store that on the wiki. For further sessions, just open the wiki and re-draw just the bit of the design that is important to the feature being discussed. Make the necessary adjusts on the white board, take another picture and stored it on the wiki again. There is much more to that, if you want to dive into agile modeling, but I will leave it to another post.

This goes way beyond transforming nouns and verbs into classes and methods. Taking the examples above, for example, client, organisation and products could be transformed in entities; route specification could be a strategy class used by a routing service (that would also need to be added to the diagram and to the ubiquitous language); sales system would be an external system that we need to integrate to, etc.

Making the code more expressive

The code should reflect all concepts exposed by the model. Classes and methods should be named according to the names defined by the domain. Associations, compositions, aggregations and sometimes even inheritances should be extracted from the model. 

Sometimes, during implementation, we realise that some of the domain concepts discussed and added to the model don't actually fit well together and some changes are necessary. When it happens, developers should discuss the problems and/or limitations with the domain experts and refactor the domain model in order to favour a more precise implementation, without ever distorting the business significance of the design. 

Ultimately, the code is the most important artefact of a software project and it needs to work efficiently. Regardless of what many experts in the subject say, code implementation will have some impact on the design. However, we need to be careful and very selective about which aspects of the code can influence design changes. As a rule, try as much as you can to never let technical frameworks limitations influence your design, but as we know, every rule has exceptions.

A common implementation problem that very often get in the way is mapping objects to databases using ORM tools. In this case, bending the model a little bit in favour of a more realistic relation among entities is not a bad thing. Just make sure that changes like that are represented in the model and understood by everyone involved.    

DOs and DON'Ts 
  • Don't try to model everything. Focus on the core of your application. We are not working on a waterfall or Unified Process project here.
  • Try to model just the key concepts of the domain problem;
  • Do not clutter your models with too much details. Keep it focused on the main responsibilities;
  • Limit your discussions and changes in the model to the business concepts (domain objects) related to the user story being discussed;
  • Don't add architectural concepts like DAOs, Actions, etc. We are not writing implementation diagrams. 
  • As your application grows, break the application into multiple models (domains), explicitly defining the context and boundaries within which a model applies. This is called Bounded Context in Domain-Driven Design.
  • Avoid thinking purely on the implementation when designing your model. Understanding and modeling the business is the most important thing here. 
Challenges of a Model-Driven Design

Model-Driven Design (MDD) is more an art than a science. It takes a lot of practice and willingness to get it going and get it right. Refactoring towards deeper insights must be seen as a positive and essential part of the project development. TDD and continuous integration are also essential for any agile and domain-driven application.

The goal of MDD is having the software expressing a deep and supple design.

Last buy not least, developers with good Object-Oriented Design skills are needed in the project. The lack of design skills could easily transform ANY software project in a total failure in the long term.

Source

Friday, 11 June 2010

The Wolf in Sheep's Clothing

In the last few years, I've noticed that the majority of the projects that I've participated roughly followed the same design. Speaking to colleagues and friends, the great majority said that they were also following the same design. It's what I call "ASD" design (Action-Service-DAO).

by Action we mean a Struts Action, Spring Controller, JSF backing bean, etc. Any class that handles an action triggered by the user interface.

We can argue that there is nothing really wrong about this approach since if we map that to a three-tier architecture, that is followed by the majority of the applications, the ASD classes would be in the right places, keeping a good isolation between the tiers.


One of the problems of this approach is when services are created randomly, with different granularities, low cohesion and with a weak representation of the business domains. Services end up being function libraries, almost like an utility class where all "functions" related to a specified "module" are grouped. When it comes to DAOs, the situation is not very different. DAOs are developed in a way that they become utility classes where sometimes we find a single DAO with all queries, inserts, delete, updates for an entire module or sometimes you find one DAO per entity. Either way, regardless what the query returns (or what it is its intention), or any rules related to updates or deletes, the methods will go to the same class. As the application grows, the code becomes something like this:


Looking at the picture above, can we really say that it is object-oriented programming? As services and DAOs don't strongly represent business concepts, the code becomes procedural. I don't want to bang on about the advantages of OOP over Procedural code. I believe that too many people have already discussed that over the past 30 years. My point here is to discuss what a "design" like that can do to an application.

Following bad examples


Unfortunately, in software development, bad examples are easier to be followed than good examples. With a non-expressive business model like that, the services get overloaded with methods. When adding new features to the application, developers need to go through all the methods of a service to see if there is already a method that does what they need. Due to the lack of cohesion and method explosion in the services, many developers will just add a new method in there, that can be very similar to existing ones, instead of re-factor the existing ones. This leads to services with confusingly similar methods and a lot of duplication. As a chain of "bad" events, the more methods a service has, the more classes depending on it the application will have. Many dependencies means that re-factoring the class would be harder, discouraging any one willing to improve the quality of the code.

Duplication does not happen just inside a single service. It's also very common to find different services with very similar methods, if not the same. In general, this is due for the lack of clarity on the responsibility of each service. According to the feature that developers are working on, they will choose a service to add (or reuse) the methods they need. If the responsibility and granularity of each service is not clear, business rules related to a certain area of the business will be spread all over the place since different developers will think that the method will belong to different classes.  


Exposing DAOs to the presentation tier

Another side effect of this approach is that since the DAOs are also "utility classes" with no business meaning (it just has an architectural meaning), some developers can easily expose the DAOs to the actions, without going through the services. Let's give more meaningful names and more details about the Action 4 and DAO 4 shown above. Let's call them ClientAction and ClientDAO.



Here the user interface needs to display a list of clients. ClientDAO has a method that returns a list of clients. In theory this may make sense. The most used argument in favour of this approach instead of adding a "ClientService" in between the ClientAction and the ClientDAO is that the service would have a method that just delegates the call to the DAO. The service layer, in this case, would be considered redundant and just pointless extra work. 


Looking at this isolated scenario, having a service in the middle really looks a bit overkill and unnecessary. However, we will probably want to do more things with a client. We will probably want to create a new client, delete or update an existing client.


Here is where the problems begin. Very rarely, we have a pure CRUD application. In general, many business rules have to be performed before or after any changes are made in the persistence tier (generally a database). For example, before inserting a new client, a credit check needs to be performed. When deleting a new client, we need to make sure that there is no outstanding balance for that client and close his or her account. We also need to archive all the orders for this client.  Whatever the application's business rules dictate. Multiple entities (or database operations) may be involved in a operation like that. We need to be able to define transaction boundaries. Besides CRUD operations, we will also have all business methods like check if client has credit, add orders to a client, etc. This is too much for a single DAO to handle. DAOs should not perform any business logic and should be hidden from the presentation tier. It should be hidden (encapsulated) even from different services. DAOs should just deal with the persistence. Business logic and transaction boundaries should be controlled by a class with business responsibilities, that in the case of this "poor" design, it would be a service.   



So even having a few methods in the service that just delegate the responsibility to a DAO, it is still a price worth paying. The service should hide (encapsulate) the details of it's implementation. A client code does not need to know how and where the data is persisted.

OK, enough of this. Exposing DAOs to the presentation tier is WRONG. Let's move on.

Moving away from the procedural code (baby steps)

The first thing to do to move away from the procedural code is to make your code more expressive and aligned to the business. We need to narrow the gap between developers and domain experts (business analysts, users, product owner or whoever knows the business rules) so we all start speaking the same language. The code must be written in a way that it expresses the business and not just architectural concepts that means nothing to the business. Actions, Services and DAOs are technical terms with no connection to any business term.

There are a few techniques that can be applied (or even combined) in order to make it possible. In future posts, I'll be talking about some of these techniques in more details. I know, it's frustrating that I will not tell you right now how to do it after have spending all this time criticising the ASD procedural design. The important thing for now is to know that what looks a good solution, since it is used by many different people and in many different projects, may not be as good as it seems.

In the meantime, there is something that we can start doing to our code in order to reduce the mess and make the necessary refactorings easier in the future.

Preparing your code for an easier transition

The following advices will help us to get our code to a point where it can be easily refactored into a more domain (business) focused approach in the future. It will still be a bit procedural but will be much more well organised and a notion of components (business components) will start to emerge.

1. Identify key concepts (domains) in your application. (e.g. Client, Order, Invoice, Trip, Itinerary, etc)
2. You can create / refactor services for each one of the key domains.
3. Try to keep a well balanced granularity for all services.
4. Avoid services for small parts of the domain. For example, do not create a service for line items. Line items should be handled by the OrderService.
5. Services are not allowed to manipulate multiple domains (with exceptions of whole-part relations - compositions)
6. DAOs are encapsulated by the services and never accessed by any other class.
7. Not every service must access a DAO, but any DAO must be accessed by one and only one service.
8. Services must delegate operations to other services if the operation is related to a different domain from the one the service is handling.
9. Services talk to each other. DAOs never talk to each other.
10. Parts (like a line item) are never exposed by a service. Service always exposes the whole (like Order). Parts are accessed via the whole. (e.g. order.getLineItems();)

With the following rules, our procedural code starts looking more like meaningful objects, with a reasonably well defined interface and some significance in terms of business.

  
As mentioned before, this is still a bit procedural but this design already solve a some of the problems discussed earlier like having more meaningful and specific services. The responsibility and boundaries of each service are more defined, making it easier for developers to look for implemented methods and re-use what it is in there, reducing duplication. 

In future posts I'll be talking about the next steps towards a more expressive and business focused code, less procedural and more object-oriented.  


If you are curious about the next steps and can't wait for my posts, have a look at:
http://domaindrivendesign.org/
http://en.wikipedia.org/wiki/Behavior_Driven_Development
http://en.wikipedia.org/wiki/Model-Driven_Architecture
http://en.wikipedia.org/wiki/OOD