As MoSKito 1.5.0 hits the maven repositorty today, so do some new features about threads.
Here's a short overview on 4 new screens added in 1.5.0 for threading issues.
NOTE: We now do have an official anotheria blog: blog.anotheria.net
This post is moved there:
http://blog.anotheria.net/?p=9
Showing posts with label java. Show all posts
Showing posts with label java. Show all posts
Friday, August 17, 2012
Thursday, August 16, 2012
Everything gets better once Oracle buys you
Well, almost.
Recently I had to retrieve my sun developer framework account, I wasn't using for 5 years now.
I was hoping I can vote on this ugly bug http://bugs.sun.com/bugdatabase/view_bug.do?bug_id=7180557 with localhost and java7 on my mac, I was writing about in previous post. Well - no. However, the process itself was funny enough, since they require you to specify both, username and email, and in my case (and I suppose many other cases) its the same, but its not that easy to figure out.
However, sun's and now oracle's mailing system works much better than the update profile dialog, I was forced to fill and submit:
Well, at least we now know what GlassFish version they are running ;-)
Recently I had to retrieve my sun developer framework account, I wasn't using for 5 years now.
I was hoping I can vote on this ugly bug http://bugs.sun.com/bugdatabase/view_bug.do?bug_id=7180557 with localhost and java7 on my mac, I was writing about in previous post. Well - no. However, the process itself was funny enough, since they require you to specify both, username and email, and in my case (and I suppose many other cases) its the same, but its not that easy to figure out.
However, sun's and now oracle's mailing system works much better than the update profile dialog, I was forced to fill and submit:
Well, at least we now know what GlassFish version they are running ;-)
Oracle Kills getLocalhost on MacOS X in Java 7
This is going to be another one of those How things that can't happen actually happen post.
I recently updated my Mac to Java 7. Pretty late though. Right after that DistributeMe services start behaving at least strange. Right after the start a service registers itself in a registry with a unique service identifier, which looks like that:
Where 192.168.1.113 being the ip adress of the machine the service runs on. This adress is used by the client, once the later wants to connect to the service.
After upgrade to Java 7 the service registered itself in the registry with the identifier:
Of course no client were able to resolve the host named unknown. I started investigating where this comes from, and well, it was in one of those sections...
I wrote a very small program to verify it:
Running it with JAVA6 and JAVA7 shows the difference, watch yourself:
Sad.
I recently updated my Mac to Java 7. Pretty late though. Right after that DistributeMe services start behaving at least strange. Right after the start a service registers itself in a registry with a unique service identifier, which looks like that:
rmi://a_b_c_FooService.qhxgttedwr@192.168.1.113:9252@20120816162707
Where 192.168.1.113 being the ip adress of the machine the service runs on. This adress is used by the client, once the later wants to connect to the service.
After upgrade to Java 7 the service registered itself in the registry with the identifier:
rmi://a_b_c_FooService.qhxgttedwr@unknown:9252@20120816162707Of course no client were able to resolve the host named unknown. I started investigating where this comes from, and well, it was in one of those sections...
private static String getHostName(){
try{
InetAddress localhost = InetAddress.getLocalHost();
String host = localhost.getHostAddress();
HashMap mappings = configuration.getMappings();
String mappedHost = mappings.get(host);
return mappedHost == null ? host : mappedHost;
}catch(UnknownHostException e){
return "unknown";
}
}
Ok, this is certainly not the best idea to return "unknown" here (one of those, can't happen anyway bug), but why the heck did InetAddress.getLocalHost() fail?I wrote a very small program to verify it:
package localhostbug;
import java.net.InetAddress;
public class PrintLocalhost {
public static void main(String[] args) throws Exception{
InetAddress localhost = InetAddress.getLocalHost();
String host = localhost.getHostAddress();
System.out.println("host: "+host);
}
}
Running it with JAVA6 and JAVA7 shows the difference, watch yourself:
$JAVA6_HOME/bin/java -cp classes localhostbug.PrintLocalhost host: 192.168.140.200and
$JAVA7_HOME/bin/java -cp classes localhostbug.PrintLocalhost Exception in thread "main" java.net.UnknownHostException: colin.speedport.ip: colin.speedport.ip: nodename nor servname provided, or not known at java.net.InetAddress.getLocalHost(InetAddress.java:1438) at localhostbug.PrintLocalhost.main(PrintLocalhost.java:7) Caused by: java.net.UnknownHostException: colin.speedport.ip: nodename nor servname provided, or not known at java.net.Inet6AddressImpl.lookupAllHostAddr(Native Method) at java.net.InetAddress$1.lookupAllHostAddr(InetAddress.java:866) at java.net.InetAddress.getAddressesFromNameService(InetAddress.java:1258) at java.net.InetAddress.getLocalHost(InetAddress.java:1434) ... 1 moreAfter some googling I found a bug in oracle bug database and a similar issue in openjdk. Seems Java and Mac is getting less and less a love story.
Sad.
Monday, June 25, 2012
MoSKito 1.4.3 released
After our first iphone app arrived in the store on friday, we also released 1.4.3 of the classic version over the weekend. As the number suggests, this release doesn't bring major breakthroughs (at least not yet), but some improvements in the UI for better user experience.
First, and the simplest of all, we adopted the default size for the on-the-fly-graphs for producers from 600x300 to 1200x600, making it four times bigger. Now you will be able to use the whole size of your monitor, and not only a small section:
Of course they are a bit smaller here, because otherwise they would kill the layout of the blog. However, if the new size is too big, or still too little for you, you can configure it now by yourself, by simply adding a file named mskwebui.json into your classpath (web-inf/classes is easiest).
The file should have the typical json config format:
Of course all ConfigureMe features like cascading environments and on-the-fly reconfiguration are supported.
This was actually a feature people have been asked for, for a long time, to allow filtering for producer names. This is especially useful, if you have a log of producers with similar names.
Here an example:
Enlarge the image for details.
Finally the accumulators overview now offers a new link:
The new version can be obtained as usual from our nexus repository.
Enjoy ;-)
Larger graphs
First, and the simplest of all, we adopted the default size for the on-the-fly-graphs for producers from 600x300 to 1200x600, making it four times bigger. Now you will be able to use the whole size of your monitor, and not only a small section:
Of course they are a bit smaller here, because otherwise they would kill the layout of the blog. However, if the new size is too big, or still too little for you, you can configure it now by yourself, by simply adding a file named mskwebui.json into your classpath (web-inf/classes is easiest).
The file should have the typical json config format:
{
"producerChartWidth": 1200,
"producerChartHeight": 600,
}
Of course all ConfigureMe features like cascading environments and on-the-fly reconfiguration are supported.
Producer filtering
This was actually a feature people have been asked for, for a long time, to allow filtering for producer names. This is especially useful, if you have a log of producers with similar names.
Here an example:
Enlarge the image for details.
Finally the accumulators overview now offers a new link:
which leads to the new:
Single Accumulator View
Single accumulator view offers a quick glance at one accumulator and provides not only the graph, but also the data behind the graph for quick analysis:The new version can be obtained as usual from our nexus repository.
Enjoy ;-)
Monday, June 18, 2012
Don't believe in availability, test for it!
Preamble
Back in the year 2004 Helmut Oertel and I were conducting a technical due diligence of a job portal for the scout group. One of the topics was backup and disaster recovery, and the questions went like this:
- What happens if your database crashes?
- We have a stand-by database server.
- Is this a hot- or cold-standby?
- It's a hot standby, but we currently switched it off.
- So it's a cold standby then?
- Aehm... probably yes...
- Have you actually ever tested it?
- .... silence ....
Now, this is something you would call epic fail. Having a spare db in theory, but never testing it, is actually worse than not having it at all. In the later case, you at least don't have the feeling of false safety.
FriendScout for example, did it better, they had two database servers in master/slave failover mode, partially coded by themselves, and they did switch the master and slave every week or so. This way they actually new, that both, master and slave, are able to work. In fact, as they had a small problem with one of the less important databases, and the master failed on Dec 24th (no kidding), they didn't detected it until Januar 3rd, because the failover was performed so smoothly, that no customer was affected.
Another good example is Parship. They have a DistributeMe based SOA and each important service is replicated in multiple instances with FailoverToNextNode failing strategy. This means that if service Foo on server 1 fails, the call is retried with server 2, and so on. This is needed in production or staging, but in local development environment it is an unnecessary overhead to start nearly 100 services on a your dev machine, so they start only one instance of each service, and make failover routing do its job. If the client (a web controller or something) is issuing a request to server X, instance 1 (based on mod-based routing of the userid for example), and that instance is not running, the failover mechanism finds a working instance of this service. In fact they achieving two goals, 1) they save resources and 2) they continuously test their failover strategies. The prove came as a buggy FailoverRouter was committed to the trunk (a combination of version incompatibilities, nearly impossible to find by unit tests). One could expect that such an error would only be detected in a real failover situation in production environment, but due to continuous failover testing it was detected 20 minutes after commit.
But lets come to the point.
Talking about DistributeMe, it has a set of built-in interceptors which are meant for availability testing.
However, the concept of interceptors is common to many (and probably most) middleware solutions, so you can easily implement the same thing with CORBA or what-ever your platform offers.
Before we start, some ontology:
you might get the feeling that I mix up service and server. I don't. In the SOA world, at least how we understand it:
a Service is a component which offers some services (methods) and which behavior is defined by contract (interface). In other word - Service is some code.
a Server is a node in the distributed system which contains (runs) one or multiple services. In other words a server is a JavaVM. The most usual situation however, is that one server runs one service.
In a mod- or roundrobin- distributed, failover enabled systems a service is running in multiple servers simultaneously.
a Servant (CORBA Slang) is the service implementation process inside the server.
So whenever a client is talking to a logical component it is talking to a service, but the physical data is transformed to a service instance in a server.
So, what are the most interesting cases you should test for. The easiest and most obvious one is surely:
Server is not there.
Server is not there is a very common case. The server could have crashed, didn't start, the real or virtual machine crashed and so on. However, this failure is easy to detect. If the server is not there, the reaction comes immediately.
To emulate this behavior the interceptor simply throws NoConnectionToServerException before the call could probably leave the stub (client side interceptor) or before the call could be delegated to the servant (server side interceptor).
Server is slow.
Server is slow is a more complicated and more dangerous situation. Generally I have encountered two reasons for otherwise healthy service to become slow. The are surely not the only possible, but they are most probable:
- Unexpected DB Problems
- Continuous Full GC cycles
An extreme version of a slow server is a never replying server, for example due to deadlock and thread starvation, but this case is pretty similar to the above in terms of consequences.
So what happens if the server is very slow. Well it depends on your architecture and your middleware. For example RMI has no thread pool limitations by default. This means that a popular server would virtually drain all the threads from your web servers, until they don't have anything left and no user request can be proceed. I have seen service instances with over 10.000 threads waiting for them, which can get pretty ugly, cause those queues will a) take time to proceed and b) overload a newly started instances of the slow service, causing the problem to repeat.
DistributeMe offers a SlowDownInterceptor which simply puts the current thread to sleep for a given amount of time, simulating slow responding service. Of course you are free to create your own, producing full CPU load on a machine instead would be an interesting options, especially regarding cross effects with other services on that machine.
So, how to protect your system against slow servers? Well it's not uncomplicated, because it's not that easy to interrupt a (hanging) thread from outside, without messing up your system state.
Right now, we have two weapons against slow servers:
- Concurrency control
- Asynchronous method calls
Concurrency control allows us to limit the number of the request that can be sent to the server/servant. It can be applied on client and/or server side. Different approaches are possible, but the simplest one is to count active connections, and not allow new connections once the limit is reached. It also shows good result to apply different limits on both sides. For example, if you have a server that can handle 100 parallel connections and 10 clients, it makes sense to apply the limit of 100 on the server side, and a limit of 20-30 on the client side. This would prevent one client from occupying all threads and letting other clients starve.
Asynchronous method calls on the other hand give you full control of acceptable call duration. You can state that if the result is not there after 2 seconds, it can be skipped at all, and provide alternative information to the user. However they aren't un-tricky, since they have internal thread pools, hard to configure and tune. But you have to die one death, and this death is probably more pleasant to die.
Flipping server
Flipping server emulates that strange unpredictable behavior that sometime happens. For example letting the server serve only 10% of the request slowly, or throw an exception from-time-to-time. Flipping at 1 percent or less can be consider as rare error, a condition that drives a lot of devops crazy worldwide. Testing for it will help you to tune your monitoring systems to be high grained enough.
How to test
Once you made your homework and prepared your system for partial failure, installed and configured all interceptors, the hard time begins - the testing.
You will need some things for successful testing:
- a reliable test which produces fair load of the system and checks all aspects of the system.
- an integration/test system to run this test against.
- deployed version of your code.
- and a lot patience.
Since you will probably never have a complete, reliable, uptodate, automatic, detailed error reporting test, I strongly recommend you, to grab the best QA guy your company has and do it together with him. If you are part of the DevOps team in your company - fine, if not, you should have one on board too, because the guys will have to solve the problems you are trying to simulate in production, and you will help them a lot if they can see and learn the symptoms.
Once you have all the guys you need, you are ready to enter the never-ending iteration:
- run the test
- detect the failures (where the system reacts badly on a service-failure)
- fix it.
repeat.
If you don't find anything to fix, try out new interceptors, manipulate data in the db, switch bytes in the packets, interrupt threads. Make availability testing part of your development process, and you will have a really responsive, robust, and almost unbreakable system you want.
AND last but not least - it's fun!
Thursday, December 22, 2011
The three most fatal bugs, ever.
There are bugs and then there are BUGs. The bugs are usually fixed and forgotten, but BUGs remain with you forever. I'd like to tell about three, that are worth telling.
The first one occurred (do bugs actually occur? pardon my english ;-) ) as I was working for FriendScout in 2005. We had a nice monitoring tool, which had a table representing the production system with a row for each webserver in the farm, and it changed the color according to the status of the server. If the server was replying it was green and if not - red. Once in August it started to switch red. One server after the next. And after the forth server it was over. The situation repeated. Once, twice, thrice... We had searched and search and had found nothing. It looked like it was a user which was doing something weird, that killed the server, but what?
Finally, Oliver, a fellow engineer, has found it, and it was a picture. But not just a picture, it was a pdf file named jpg. A lady wanted to upload a picture, which was a pdf. Pdf's weren't supported mime types (only jpegs and gifs were back then), so the upload servlet rejected the file. But the lady was (or thought she was) smarter than us and renamed the file to jpg. Now it passed the mime-check all-right and was passed over to the image processing software, a more or less standard tool which was called imagemagick. But the tool was also smarter than us and ignored the file name, extension and mime-type all together, instead it looked into the file and detected it was a pdf. After this glorious discovery it tried to call a pdf processing tool - ghostscript. But since we never actually wanted to process or even accept pdfs, ghostscript wasn't installed on the servers and the attempt to start it crashed the imagemagick lib. Since the code was native it took the whole JVM with it. Ouch.
The second bug is the power of the debug log. Back in prehistoric version of ASG I was hacking a very first version of this site. It worked absolutely great, the customer loved it and everything. After some time the customer called me and told me that the site was getting slower each day. I looked at my local copy - fast as wind. Live installation - everything ok, but adding new items (machines back then) lasts about a minute. I was searching for the bug for three days, as the customer called again, and said that creating new cms items is now about 2 minutes each, and that the only thing he did was adding 100 items. Now it was at least something. I double-checked the logs - nothing. I reviewed the code - nothing. I had no good profiler (yes, if I had MoSKito back than, I would have found it faster), so I started to add time measurements everywhere, and, after a long hunt, I finally found the line - it was a log.debug statement.
The service that was responsible for storing the items in the cms had one small innocent line: log.debug(cache). Since debug output was off, the log call had no visible effect, hence I had nothing in logs, but the cache got bigger and bigger with each added item, and the effort to execute its toString method, which printed out all contained elements, was growing constantly. From 0 seconds in the unit test to 180 in production. That teached me to use if log.isDebugEnabled()... At least something ;-)
And this was the second bug. Now for my personal all time favorite, super-bug we have to return back in text and time. It also happened at FriendScout but 2003, before they hired me (and this may well have been the reason why ;-)) The platform was pretty unstable and run by people, who were thinking more like admins as like developers. And since the admin usually doesn't understand, why something is broken, he has a standard method to fix things -> the infamous alt-ctrl-del button. In this case it was a little bit more subtile that that, so some smart (admin) guy wrote a script that was watching the log files and searching for the keyword 'FATAL' in it. If it would find an occurrence of FATAL, it would assume a bad-bad-bad error happened and trigger a full application restart. And there were some of the restarts in the year until they had to review this strategy. The reason for the review was a customer who called the customer desk and asked: "why does the system ALWAYS crash when I log in?"
The answer was easy, it was her login name. It was femme_fatale.
So what about you? Do you have funny bugs worth mentioning? Tell me please! ;-)
The first one occurred (do bugs actually occur? pardon my english ;-) ) as I was working for FriendScout in 2005. We had a nice monitoring tool, which had a table representing the production system with a row for each webserver in the farm, and it changed the color according to the status of the server. If the server was replying it was green and if not - red. Once in August it started to switch red. One server after the next. And after the forth server it was over. The situation repeated. Once, twice, thrice... We had searched and search and had found nothing. It looked like it was a user which was doing something weird, that killed the server, but what?
Finally, Oliver, a fellow engineer, has found it, and it was a picture. But not just a picture, it was a pdf file named jpg. A lady wanted to upload a picture, which was a pdf. Pdf's weren't supported mime types (only jpegs and gifs were back then), so the upload servlet rejected the file. But the lady was (or thought she was) smarter than us and renamed the file to jpg. Now it passed the mime-check all-right and was passed over to the image processing software, a more or less standard tool which was called imagemagick. But the tool was also smarter than us and ignored the file name, extension and mime-type all together, instead it looked into the file and detected it was a pdf. After this glorious discovery it tried to call a pdf processing tool - ghostscript. But since we never actually wanted to process or even accept pdfs, ghostscript wasn't installed on the servers and the attempt to start it crashed the imagemagick lib. Since the code was native it took the whole JVM with it. Ouch.
The second bug is the power of the debug log. Back in prehistoric version of ASG I was hacking a very first version of this site. It worked absolutely great, the customer loved it and everything. After some time the customer called me and told me that the site was getting slower each day. I looked at my local copy - fast as wind. Live installation - everything ok, but adding new items (machines back then) lasts about a minute. I was searching for the bug for three days, as the customer called again, and said that creating new cms items is now about 2 minutes each, and that the only thing he did was adding 100 items. Now it was at least something. I double-checked the logs - nothing. I reviewed the code - nothing. I had no good profiler (yes, if I had MoSKito back than, I would have found it faster), so I started to add time measurements everywhere, and, after a long hunt, I finally found the line - it was a log.debug statement.
The service that was responsible for storing the items in the cms had one small innocent line: log.debug(cache). Since debug output was off, the log call had no visible effect, hence I had nothing in logs, but the cache got bigger and bigger with each added item, and the effort to execute its toString method, which printed out all contained elements, was growing constantly. From 0 seconds in the unit test to 180 in production. That teached me to use if log.isDebugEnabled()... At least something ;-)
And this was the second bug. Now for my personal all time favorite, super-bug we have to return back in text and time. It also happened at FriendScout but 2003, before they hired me (and this may well have been the reason why ;-)) The platform was pretty unstable and run by people, who were thinking more like admins as like developers. And since the admin usually doesn't understand, why something is broken, he has a standard method to fix things -> the infamous alt-ctrl-del button. In this case it was a little bit more subtile that that, so some smart (admin) guy wrote a script that was watching the log files and searching for the keyword 'FATAL' in it. If it would find an occurrence of FATAL, it would assume a bad-bad-bad error happened and trigger a full application restart. And there were some of the restarts in the year until they had to review this strategy. The reason for the review was a customer who called the customer desk and asked: "why does the system ALWAYS crash when I log in?"
The answer was easy, it was her login name. It was femme_fatale.
So what about you? Do you have funny bugs worth mentioning? Tell me please! ;-)
Using asynchronous method calls to mine gold faster.
Disclaimer:
First of all let me say, that we are not going to mine any real gold here. However, gold mining seemed as a good example, since the wins through asynchronous communication can be as valuable.
Second, code examples in this post rely on a specific programming language (java) and a specific framework (will be named at the end) but can well be used with any other language or framework. Talking about frameworks, asynchronous communication is all but trivial and its of great value to have a synchronously programmable framework which hides the details from the developer.
Lets go back a century or two and dig for some gold. For that we have developed a special service which offers two methods:
You probably noticed the @DistributeMe annotation on the top of the class. More on it later.
The first method searches for gold at a location. When we call it, it picks a random location and starts digging. It digs for several seconds (1 to 10) making a meter per second until it finds something. And than this something can be gold or clay. Two gold miners are using this service, the SynchGoldSearcher and the AsynchGoldSearcher. Both have 60 seconds to find possible locations for a gold mine.
The SynchGoldSearcher is pretty straight forward, it just digs until it hits something. Just like the classical service call does. So please allow me to introduce Player1:
There is no real magic in the above code, so I'll skip the comments. Now the second player looks pretty similar except for a small detail:
First it asks for an AsynchRemote instead of a 'normal' Remote. Second, a CallTimeoutedException comes into play.
Remember the annotation @DistributeMe on top of the class? It belong to the distributeme framework and tells it to generate RMI code for my interface. Among other stuff it tells the framework that I want to have an asynchronous stub and that the default timeout for calls over this stub are 2500 ms. This means that any call that lasts longer than 2500 ms will be aborted and a CallTimeoutedException will be thrown.
Now lets start both searchers:
The synchronous searcher provides this, shortened output:
The asynchronous searcher provides this, also shortened output:
The asynchronous approach seems to have been more successful? I admit, I run the test multiple times to achieve the results I wanted, but the number of found mining locations are not that important to the showcase as the number of the attempts. The client doesn't really know, how long the server will need to answer the request. In fact there are situation when an RMI call will hang nearly forever on the server side, and you'll be unable to do anything about it. With asynchronous approach the calling thread is detached from the call and from the server side processing, therefore allowing you (or the framework) interrupt it at any time and return.
When can it be useful? There are many scenarios, but the most important is the quality of service to the user. With this approach you can guarantee that the site will be responsive and answer within acceptable limits (even if some times this answer will be: I can't get this information right now).
But wait, we are just getting warm. Most of the web applications and especially portals have pages where they need to combine multiple pieces of information. Often the process of retrieving the information is long and retrieving the information sequentially sums up many long retrievals. Ironically, often enough the resources in question are not competing (for example two database servers) and we could speed up the retrieval by getting the information in parallel, but its just such a hassle to do all the concurrency programming. Thankfully its not ;-)
To demonstrate this, here is another example. After we got some productive mines in previous section we are now wash mined gold. To achieve this we call the method washGold which washes the gold for given amount of time, producing one clump in each second. Again, we have a synchronous and an asynchronous versions, first the synchronous one:
Again, we compare them against each other, first the synchronous one:
and the asynchronous one:
Now this is a difference! 5 Times faster? Well, of course it was an easy win, since you as a smart reader already noticed that I'm starting the 5 calls in parallel. But the actually amazing thing is that its achieved with minimal code overhead!
If we take a look at the source code, we see that we only need few additional lines to change the behavior of sequential code and we are programming sequentially even we run concurrently:
First we create a new collector and tell it that we are going to call 5 methods (calls is 5):
Than we call the methods asynchronously via the auto generated asynchronous interface. Each calls returns immediately.
Now we tell the collector that we are ready calling and want to wait for max 11 seconds for results.
And finally we have to collect the results.
And even if we are calling 5 times to the same service, there are no physical limitations to the amount of called services or methods.
So where would you use it? Example one: you have a portal with a welcome page which gets some information from multiple modules in the system: new messages, favorites online, the weather and so on. Instead of waiting for each of the submodules you only need to wait for the slowest one. And if its too slow, you can always abort it by setting the max call duration via timeout or similar.
Another example: you have to perform a complicated and long lasting calculation, you split it in multiple blocks and let multiple nodes calculate parts of it. It could save your scaleability problems, of course only if your problem is generally dividable in multiple subtasks, but most are.
To round it up, the above can very well be achieved by programming it manually and will help you a lot in running a high trafficked site or a non-trivial calculation. You don't need to have a specific framework, but if you want to save yourself the hassle you can use ours ;-)
More on DistributeMe.
Source code of the examples.
Exact instructions how to run the tests on your local machine will be provided later... if requested ;-)
First of all let me say, that we are not going to mine any real gold here. However, gold mining seemed as a good example, since the wins through asynchronous communication can be as valuable.
Second, code examples in this post rely on a specific programming language (java) and a specific framework (will be named at the end) but can well be used with any other language or framework. Talking about frameworks, asynchronous communication is all but trivial and its of great value to have a synchronously programmable framework which hides the details from the developer.
Lets go back a century or two and dig for some gold. For that we have developed a special service which offers two methods:
@DistributeMe(asynchSupport=true, asynchCallTimeout=2500)
public interface GoldMinerService extends Service{
/**
* Searches a random location for gold. Can last up to 10 seconds.
* Returns if anything was found.
* @return
*/
boolean searchForGold();
/**
* Washes gold for a given duration. Returns the amount of washed clumps.
* @param duration
* @return
*/
int washGold(long duration);
}
You probably noticed the @DistributeMe annotation on the top of the class. More on it later.
The first method searches for gold at a location. When we call it, it picks a random location and starts digging. It digs for several seconds (1 to 10) making a meter per second until it finds something. And than this something can be gold or clay. Two gold miners are using this service, the SynchGoldSearcher and the AsynchGoldSearcher. Both have 60 seconds to find possible locations for a gold mine.
The SynchGoldSearcher is pretty straight forward, it just digs until it hits something. Just like the classical service call does. So please allow me to introduce Player1:
public class SynchGoldSearcher {
public static void main(String[] args) throws Exception{
GoldMinerService service = ServiceLocator.getRemote(GoldMinerService.class);
long start = System.currentTimeMillis();
int searchTime = 60;
System.out.println("Searching for gold for "+searchTime+" seconds");
long endTime = start + 1000L*searchTime;
long now;
int foundGold = 0;
int attempts = 0;
while ((now = System.currentTimeMillis())<endTime){
System.out.println("Attempt "+(++attempts));
if (service.searchForGold()){
foundGold++;
System.out.println("Found gold!");
}else{
System.out.println("Nothing here...");
}
}
System.out.println("Found "+foundGold+" gold in "+attempts+" attempts and "+(now - start)/1000+" seconds.");
}
}
There is no real magic in the above code, so I'll skip the comments. Now the second player looks pretty similar except for a small detail:
public class AsynchGoldSearcher {
public static void main(String[] args) throws Exception{
GoldMinerService service = ServiceLocator.getAsynchRemote(GoldMinerService.class);
long start = System.currentTimeMillis();
int searchTime = 60;
System.out.println("Searching for gold for "+searchTime+" seconds");
long endTime = start + 1000L*searchTime;
long now;
int foundGold = 0;
int attempts = 0;
while ((now = System.currentTimeMillis())<endTime){
System.out.println("Attempt "+(++attempts));
try{
if (service.searchForGold()){
foundGold++;
System.out.println("Found gold!");
}else{
System.out.println("Nothing here...");
}
}catch(CallTimeoutedException timeoutException){
System.out.println("too deep, aborted...");
}
}
System.out.println("Found "+foundGold+" gold in "+attempts+" attempts and "+(now - start)/1000+" seconds.");
((AsynchStub)service).shutdown();
}
}
First it asks for an AsynchRemote instead of a 'normal' Remote. Second, a CallTimeoutedException comes into play.
Remember the annotation @DistributeMe on top of the class? It belong to the distributeme framework and tells it to generate RMI code for my interface. Among other stuff it tells the framework that I want to have an asynchronous stub and that the default timeout for calls over this stub are 2500 ms. This means that any call that lasts longer than 2500 ms will be aborted and a CallTimeoutedException will be thrown.
Now lets start both searchers:
The synchronous searcher provides this, shortened output:
Searching for gold for 60 seconds
Attempt 1
Nothing here...
Attempt 2
Nothing here...
...Attempt 7
Found gold!
... Attempt 10
Nothing here...
Found 1 gold in 10 attempts and 65 seconds.
The asynchronous searcher provides this, also shortened output:
Searching for gold for 60 seconds
Attempt 1
too deep, aborted...
Attempt 2
Nothing here...
Attempt 3
too deep, aborted...
Attempt 4
too deep, aborted...
Attempt 5
Found gold!
Attempt 6
too deep, aborted...
Attempt 7
Found gold!
...
Attempt 28
too deep, aborted...
Found 6 gold in 28 attempts and 62 seconds.
The asynchronous approach seems to have been more successful? I admit, I run the test multiple times to achieve the results I wanted, but the number of found mining locations are not that important to the showcase as the number of the attempts. The client doesn't really know, how long the server will need to answer the request. In fact there are situation when an RMI call will hang nearly forever on the server side, and you'll be unable to do anything about it. With asynchronous approach the calling thread is detached from the call and from the server side processing, therefore allowing you (or the framework) interrupt it at any time and return.
When can it be useful? There are many scenarios, but the most important is the quality of service to the user. With this approach you can guarantee that the site will be responsive and answer within acceptable limits (even if some times this answer will be: I can't get this information right now).
But wait, we are just getting warm. Most of the web applications and especially portals have pages where they need to combine multiple pieces of information. Often the process of retrieving the information is long and retrieving the information sequentially sums up many long retrievals. Ironically, often enough the resources in question are not competing (for example two database servers) and we could speed up the retrieval by getting the information in parallel, but its just such a hassle to do all the concurrency programming. Thankfully its not ;-)
To demonstrate this, here is another example. After we got some productive mines in previous section we are now wash mined gold. To achieve this we call the method washGold which washes the gold for given amount of time, producing one clump in each second. Again, we have a synchronous and an asynchronous versions, first the synchronous one:
public class SynchGoldWasher {
public static void main(String[] args) throws Exception{
GoldMinerService service = ServiceLocator.getRemote(GoldMinerService.class);
int washTime = 10;
System.out.println("Washing gold for "+washTime+" seconds.");
long start = System.currentTimeMillis();
int washed = service.washGold(washTime*1000L);
long duration = System.currentTimeMillis() - start;
System.out.println("Washed "+washed+" gold clumps in "+duration+" ms.");
}
}
and now the asynchronous one:
public class AsynchGoldWasher {
public static void main(String[] args) throws Exception{
GoldMinerService service = ServiceLocator.getAsynchRemote(GoldMinerService.class);
AsynchGoldMinerService asynchService = (AsynchGoldMinerService)service;
int washTime = 10;
int calls = 5;
System.out.println("Washing gold for "+washTime+" seconds in "+calls+" calls.");
long start = System.currentTimeMillis();
MultiCallCollector collector = new MultiCallCollector(calls);
for (int i=0; i<calls; i++){
asynchService.asynchWashGold(washTime*1000L, collector.createSubCallHandler(""+i));
}
collector.waitForResults(11000);
int washed = 0;
for (int i=0; i<calls; i++){
washed += (Integer)collector.getReturnValue(""+i);
}
long duration = System.currentTimeMillis() - start;
System.out.println("Washed "+washed+" gold clumps in "+duration+" ms.");
asynchService.shutdown();
}
}
Again, we compare them against each other, first the synchronous one:
Washing gold for 10 seconds.
Washed 10 gold clumps in 10193 ms.
and the asynchronous one:
Washing gold for 10 seconds in 5 calls.
Washed 50 gold clumps in 10200 ms.
Now this is a difference! 5 Times faster? Well, of course it was an easy win, since you as a smart reader already noticed that I'm starting the 5 calls in parallel. But the actually amazing thing is that its achieved with minimal code overhead!
If we take a look at the source code, we see that we only need few additional lines to change the behavior of sequential code and we are programming sequentially even we run concurrently:
First we create a new collector and tell it that we are going to call 5 methods (calls is 5):
MultiCallCollector collector = new MultiCallCollector(calls);
Than we call the methods asynchronously via the auto generated asynchronous interface. Each calls returns immediately.
for (int i=0; i<calls; i++){
asynchService.asynchWashGold(washTime*1000L, collector.createSubCallHandler(""+i));
}
Now we tell the collector that we are ready calling and want to wait for max 11 seconds for results.
collector.waitForResults(11000);
And finally we have to collect the results.
int washed = 0;
for (int i=0; i<calls; i++){
washed += (Integer)collector.getReturnValue(""+i);
}
And even if we are calling 5 times to the same service, there are no physical limitations to the amount of called services or methods.
So where would you use it? Example one: you have a portal with a welcome page which gets some information from multiple modules in the system: new messages, favorites online, the weather and so on. Instead of waiting for each of the submodules you only need to wait for the slowest one. And if its too slow, you can always abort it by setting the max call duration via timeout or similar.
Another example: you have to perform a complicated and long lasting calculation, you split it in multiple blocks and let multiple nodes calculate parts of it. It could save your scaleability problems, of course only if your problem is generally dividable in multiple subtasks, but most are.
To round it up, the above can very well be achieved by programming it manually and will help you a lot in running a high trafficked site or a non-trivial calculation. You don't need to have a specific framework, but if you want to save yourself the hassle you can use ours ;-)
More on DistributeMe.
Source code of the examples.
Exact instructions how to run the tests on your local machine will be provided later... if requested ;-)
Subscribe to:
Posts (Atom)






