Showing posts with label asynchronous. Show all posts
Showing posts with label asynchronous. Show all posts

Wednesday, September 6, 2017

Promises, promises: a case for synchronous execution

On occasion I find myself counted as a resource for both front-end development and server-side development, particularly when the server-side development is in nodejs. I suppose that’s one of the advantages of having written in JavaScript since before it was JavaScript. One of the difficulties I’ve found in that endeavor, however, is when engineers who are less experienced with JavaScript start using Google to answer their questions and those answers duplicate some very poor habits. This quickly becomes the case when one of those poor habits - the overuse, or inappropriate use, of asynchronous functions like Promises, for example - has become very popular.

Not wishing to be misunderstood, I should perhaps rewind a bit and explain. In front-end development, we learned very early that JavaScript is single-threaded, and as such, would block processes. As front-end engineers concerned about performance, this was drilled into our heads at every opportunity. This was especially true for those of us writing interfaces where performance was more than just a ‘nice to have’ feature - interfaces where delays of a few seconds in a process would reduce revenue by tens of millions of dollars.

In those days, we did everything we knew to reduce blocking processes and even developed interfaces that integrated custom events and event handlers (before we actually called it the pub-sub model). When server-side JavaScript came along - for real this time- it made sense to a lot of people to use the same sort of models used in developing for user-agents because by that time we'd had years of to learn about JavaScript. What did all of our experience with JavaScript in a user-agent teach us? All of our experience taught us two things that stood out above all others: JavaScript is single-threaded and blocking processes is not nice. The folks who wrote about nodejs even wrote extensively about how the functions you write shouldn’t block processes and the ‘best practices’ discouraged use of the *sync methods - like readFileSync or writeFileSync - in the nodejs core modules. All this came about because (say it with me) JavaScript is single-threaded and that was a significant problem in code running in a user-agent.

There is, however, a dirty little secret that few in this new world want to acknowledge - some processes should be blocked.

We’ve heard it said that if you have a hammer everything looks like a nail, and that's certainly been true in this instance. However, coding a server-side process is categorically different than coding for a user-agent. In a user-agent, blocking a process impedes a person's action - action that is often randomly ordered rather than sequential. Blocking action increases the cognitive load of the process, which in turn increases the amount of time required to complete the process the user is trying to complete, and longer processes cost more. As an aside here, I should point out that even in user-agent we recognize that there are times we should or even must block processes.

So, how does this affect current practices that encompass server-side development? The focus on asynchronous processes - a large part of front-end development for very good reasons - has led to the proliferation of Promise objects. On the server side, there has been a corresponding decrease in the implementation of synchronous methods alongside it.

At this point you may be asking why, or even if, that's bad, because JavaScript is single-threaded and blocking processes is not nice, and the short answer is "yes". This quickly becomes clear if we consider a simple case in which process logging and auditing becomes critical - a process for which we must have those services in place before continuing or exiting the process. There are still deeper reasons, however.


The answer is also yes because - and I know this will be a stretch for some readers - there are a lot of differences between server-side processes and user-agent processes. As a clue to what some of those differences might be, one of the descriptions includes the word user. Even if we set aside how a user injects asynchronous behavior into a process and how that alters how we think about designing a process, there are still sufficient answers to why the proliferation of Promises in server-side development is disadvantageous.


Good code is well written...and, I've said this before (and even written it) but it bears repeating - good writing has clarity. Writing what should be a synchronous process as a then clause on a Promise masks that the process is synchronous. That masking of reality is not only not clear, it is the opposite of clear. While one might posit that writing can be good even if it lacks clarity, it certainly is not good when its true nature is cloaked in obscurity. Additionally, even though a "Promise chain" with a series of then clauses may be organized, it certainly does not have greater clarity than code that is explicit about its true, synchronous nature.

Good code is as simple as possible. We must, for a number of reasons, consider complexity the antithesis of good code - not the least of these reasons is that complex systems fail in complex ways. Coding a synchronous process as a Promise chain introduces unnecessary complexity that goes beyond a lack of clarity. Further, when Engineer Adam writes a process as a Promise and Engineer Beatrice has to use that in a synchronous process she will have to introduce await (as a module dependency until it becomes part of the native implementation), further increasing the complexity.

While the journey is interesting - as the "how did we get here" questions often are - we must move beyond that as the question in the mind of senior engineers becomes "how do we best address this situation".

The first step in addressing any problem is admitting there is one. We, collectively, must recognize that just as there are legitimate uses of asynchronous processes, there are legitimate uses of asynchronous processes. As long as we insist on the vilification of a synchronous process, we will never move forward.

Admitting there is a problem is merely the first step, however. Beyond that we, as engineers, must commit to building interfaces that allow synchronous use as well as asynchronous use. We must also commit to evaluating the execution needs of the application we are building to determine the appropriate course of action, and then following that course of action even in the face of the hue and cry of those who would insist we follow an asynchronous path.

As stated earlier, there are times in which asynchronous execution is the appropriate solution, but we should not adhere to that simply because nodejs evangelists or anyone else says we should.

Happy coding.


Friday, August 26, 2016

Free Rum

People have asked me what has helped me debug nasty, not-so-easily-reproduced bugs and nearly always the answer is "exhaustive testing"...but what about when it isn't. Ok, most of the time it really is exhaustive testing. Of course, I make it a habit to write code to handle all different data types even when I'm 99.999% sure it will always only be one data type, because when you have 100M+ users using your code daily you'd be surprised how often 0.001% is.

So, how do we debug issues that seem to only occur 0.001% of the time? Obviously we're not talking about those use cases that fall into within the "normal" range, we're talking about outliers - or what I refer to as the oddball case. Here we're talking about those times when all our unit tests pass and still our code fails in the "real world" where there are a few cases - extreme corner cases, admittedly - when users are experiencing a "sub-par visit". These instances can easily bore a hole into our confidence (and ego...it's a little bit ego), leaving us scratching our head in confusion and frustration. What do we do then?

As with any good pirate, the answer is RUM, and lots of it!

Go ahead, start making your own pirate references and talk like a pirate...I'll wait...and I'm pretty sure I can hear you singing....
"Fifteen men on the dead man's chest —
...Yo-ho-ho, and a bottle of rum!
Drink and the devil had done for the rest —
...Yo-ho-ho, and a bottle of rum!"
Robert Louis Stevenson, Treasure Island

Of course in our case RUM does not refer to the pirate's beverage of choice but refers, instead, to Real-time (or sometimes Real) User Monitoring. The emphasis here, obviously, is on user monitoring, not testing or assuming what might be happening, and it's different than synthetic user monitoring, which is a simulation (and typically done as part of a comprehensive testing plan) because it uses real people in real-world situations.

This last point - that it uses real people in real-world situations - cannot be stressed enough. Why? Because, in the specific situation this post addresses, our software engineering has begun to move away from how people have evolved. People typically solve problems in a linear manner - it's how we've evolved. C came about because A and B happened. We make our selections in a store, go to the checkout, and pay. We don't suddenly jump out of line to go to the bank and apply for a credit card and expect to return to our place in the line with our basket full of our selections. For a very long time, our software matched this model exactly, or very nearly so. Oh, we might have a decision point where we would loop back into a process, but it was still a linear process. Most of the time, this works well, but as anyone who's stood in a queue hoping to order food only to find the customer ahead of them hasn't quite decided what they want knows, sometimes there are problems with synchronous, single-thread experiences.

We tried to resolve the synchronous, single-thread problem by adding threads. This often works alright, for the most part, but as we discovered, multiple blocked threads are not really any more productive. The answer, then, was an asynchronous (non-blocking) approach. Now, here we are, years later often suffering in callback hell and struggling to debug software because our linear, synchronous experience no longer applies to development of asynchronous systems...and as was mentioned earlier, that's why RUM enters our toolkit.

There are, of course, many ways to implement (or pour) RUM. If you're running any one of several traditional web servers, IIS or Apache, for example, Splunk is a good choice. I've written a proxy server for a Splunk app and reviewed a couple Splunk references - Splunk Developer's Guide and Learning Splunk Web Framework - so I'm not without respect; however, the downside of Splunk is that it can become expensive. For this reason alone, even if your operation uses Splunk in the production environment you may choose to forego it as a support for research and development.

As a no-cost alternative, however, you can pour your own RUM if you're using a node.js server. That's right, I said it - no-cost and RUM together - FREE RUM - (part of) every pirate's dream. Since it's relatively easy to do, and as long as you don't do something really boneheaded, you can build it securely, hoist the Jolly Roger, talk like a pirate, and follow the map below. Be warned, I'm not pointing out all the dangers (like how you can expose private or confidential information) - there are a few (fairly obvious) pitfalls, but this should be enough to get you in the general area of the treasure you seek (and I should add that you can find a more complete example/prototype in my rum github repo - https://github.com/hrobertking/rum).

First, you'll need to assemble a crew and put them on a ship. Do that by installing the socket.io module, adding it to your server module, and binding it to your httpServer instance.


Server-side JavaScript (Socket instantiation)

var server = http.createServer(handler),     io = require('socket.io')(server);

You should note that the server need not be the native node.js server, it can be an extension of the node.js httpServer object - e.g., an Express instance.

Now that you have a crew and ship, decide when you'll raise the Jolly Roger - do that by emitting an event through the socket and passing an object into the event, e.g., io.emit(event_type, event_object);.

The event type is a string literal so it can be named (nearly) anything and you can emit different events at different times, as in the example below.


Server-side JavaScript (Event Handlers)

var msg = {   id: unique_id(),   req: request,   res: response };   msg.req.on('end', function() {     msg.action = 'request';     io.emit('message', msg);   });   msg.res.on('finish', function() {     msg.action = 'response';     io.emit('message', msg);       if (msg.res.statusCode !== 200) {       io.emit('http error', msg);     }   });


Now put the spyglass to your eye and scan the horizon. Your spyglass is going to be a static page that uses the socket.io client script (https://cdn.socket.io/socket.io-1.3.5.js) and instantiates a socket. That socket will then be monitored and your event handlers will run when the event comes over the socket.


Your 'spyglass' document

<!DOCTYPE html> <html lang="en">   <head>     <meta http-equiv="Content-Type" content="text/html; charset=utf-8">     <style type="text/css">       .request { background-color:red; }       .response { background-color:yellow; }       .request.response { background-color:green; }     </style>   </head>   <body>     <script src="https://cdn.socket.io/socket.io-1.3.5.js"></script>     <script>       var socket = io();             socket.on('message', function(obj) {           var node = document.getElementById(obj.id), cls;           if (!node) {             node = document.createElement('div');             node.id = obj.id;             document.body.appendChild(node);           }           cls = node.className.split(' ');           cls.push(obj.action);           node.className = cls.join(' ');           node.innerHTML = '<p>' + obj.req.url + '</p>';         });             socket.on('http error', function(obj) {           var node = document.getElementById(obj.id);           node.className += 'http-error';           node.innerHTML += '<p>' + obj.res.statusCode + '</p>';         });     </script>   </body> </html>


Now, as you watch the 'spyglass' document in your browser, when the 'io' events are fired on the node.js httpServer they will be handled in the handlers specified in the spyglass.

Now you have insight into the important pieces of code, not as you test them (because you can get that elsewhere, like the Chrome dev tools), but as others use them, and it's debugging in those real-world situations that takes our development to the next level in our drive for results. Now, go get some pirate booty.

Happy coding.

Sunday, April 5, 2015

Logging Long-running Processes

In an earlier post I talked about getting the amount of time required for long-running processes - a simple approach to determining how much time remains in this batch sort of thing. That, of course, started me thinking...

What if I have a long-running process that is one in a series of long-running processes - how can I track that?

Let's say, for example, you're going to copy several repositories from one location to another, so you set up a small script to copy them - you might have several lines, each something like "cp <source> <destination>". What if, you wanted to track the process a little more closely and wanted other people or processes to also be able to see the progress?

First, you probably don't want to use an OS command - that won't give you enough information, so you set up a small script that will copy from the source to the destination, and since this is a long-running process you want to make sure you're writing information to a log that it's started (and you're going to include something that your other scripts can use to determine how much time remains). Now your (bash) script - the one you're going to use in your batch-process script - is looking something like this...

  1. SOURCE="$1"
  2. DESTINATION="$2"
  3. LOG_FILE=`basename $0`
  4. LOG_FILE="${LOG_FILE/.sh/.log}"
  5. FILES=`ls -l $SOURCE | egrep '^-' | wc -l`
  6. STARTED=`date "+%FT%T"`
  7. LOG_ENTRY="$SOURCE\t$DESTINATION\t$STARTED\t$FILES"
  8. echo -e "$LOG_ENTRY" |tee -a $LOG_FILE
  9. cp $SOURCE $DESTINATION

Now your script will send output showing the source and destination as well as the number of files that will be copied to a log that other people and processes can read. This way when your batch file runs to copy all the repositories, the start of each copy will be recorded. That's something, at least.

Let's go a little further, though and say that you want to also record, in the same log file, not only the time the process started, but when it ended also. In that case, we're going to change the format of the log - let's insert a column between the STARTED column and the FILES column, and we need to add lines after the copy command to record the time the process ended and modify the log entry so that we're not parsing multiple lines for a single action. This gives us a script that looks like...

  1. SOURCE="$1"
  2. DESTINATION="$2"
  3. LOG_FILE=`basename $0`
  4. LOG_FILE="${LOG_FILE/.sh/.log}"
  5. FILES=`ls -l $SOURCE | egrep '^-' | wc -l`
  6. STARTED=`date "+%FT%T"`
  7. LOG_ENTRY="$SOURCE\t$DESTINATION\t$STARTED\t\t$FILES"
  8. echo -e "$LOG_ENTRY" |tee -a $LOG_FILE
  9. cp $SOURCE $DESTINATION
  10. ENDED=`date "+%FT%T"`
  11. NEW_ENTRY="$SOURCE\t$DESTINATION\t$STARTED\t$ENDED\t$FILES"
  12. sed -i.bak "s/$LOG_ENTRY/$NEW_ENTRY/" $LOG_FILE

A few notes: (1) notice line 10 has a delimitor for the blank column - that's necessary to simplify the log reader so we don't have to redefine the fourth column; (2) notice the ".bak" after the inline option for sed - that's necessary to ensure that some OSs, which require it, won't fail; and (3) if LOG_ENTRY or NEW_ENTRY contain a "/", be sure to escape them (or change sed's delimiter, which is probably easier).

Now, when your batch-process script invokes your process script, it will update the log with the SOURCE, DESTINATION, STARTED, and FILES entries when the process begins, and will update that log entry, inserting the ENDED value when the process ends.

You can use this approach to not only fine-tune your reporting and estimates, but this approach will also indicate, at a glance, which process(es) are still running...which could be even more value if the calls are asynchronous - but asynchronous calls are another topic altogether.

Happy coding.

Wednesday, March 4, 2015

Time is a river of events, and strong is its current

Time is a river of events, and strong is its current;
no sooner is a thing brought to sight than it is swept by
and another takes its place, and this too will be swept away.
~ Marcus Aurelias ~
It is doubtful that one can survive in today's software engineering environment without an understanding of asynchronous, event-driven development, and that's especially true for web developers. It's seems a little difficult to believe, but we've had AJAX for more than 10 years, and for the last 6 years or so we've had a very powerful asynchronous server-side tool in nodejs.

Is being asynchronous all it's cracked up to be, however? You may be asking yourself, if I'm just building a tool that does one thing, and I know the order in which the tasks need to be done, why not just use synchronous code...and that is a really good question. Let's take a hypothetical journey and see how easily we navigate this river of time using a real-time approach.

Let's say I am building a website. To keep costs low and security high, I'm going open-source and using Apache as my web server and I'm using Apache Sling™ to manage my content. Now I can easily build a single website and get it working without a problem; however, my use case requires that I set up thousands of 'child' sites - one for each of my local offices. Since this is a hypothetical, we don't need to bother with what sort of product is being offered - it could be an educational institution with satellite offices or a religious institution, for example - we just need to serve each local 'office' as it presents different hours, a different location, and offers slightly different services.

I can roll this whole parent-child monster out by creating the right architecture - a master template with child templates and specialized content. Doing this manually is feasible on a very small scale, but quickly becomes something that cannot be managed easily. One possible solution to this dilemma is for each office to be responsible for building their own custom site using the master site we set up; however, a significant amount of education and pre-configuration would be required to make this option a reality, as each office will need either dedicated attention or an in-house 'expert'.

Luckily for us, Apache Sling™ is itself built on a ReSTful framework, so all we need to do is gather the data for each of the offices and then use curl or some other utility to generate the correct HTTP commands to create the site, pass in the data to update all the special values, and voilĂ  - we build a website in less time than a human can even react to a signal. Easy, right? Not so fast.

When a person is using the interface to create the website, then going through and changing the content, there is a lower limit on the time it takes to perform those activities. Assuming the web server responded instantaneously (which we know is a fantasy itself), there still would be a response and hundreds of milliseconds between requests. What happens when a tool, like a team of paddlers in a raft, issue the HTTP requests too quickly, purely asynchronously without waiting for a response? There are a few possibilities - the web server starts seeing your attempts at creating and updating content as a DoS attack and responds accordingly, or it doesn't respond distinctly to a DoS attack and is overwhelmed, or we have a race condition between the requests. All of these possibilities leave you without a site at the end of the process, and possibly with a non-functional web server. As a Dungeon Master might say, "the stronger and faster among you paddled fiercely, spinning your boat in circles as the strong current of time washed away your boat".

What is needed in this case, is a tour guide who can tell those who are propelling the boat when to paddle and where. Let me demonstrate this idea by addressing the hypothetical situation by building a nodejs application that will create the website, then update the content, and finally, declare success. In order to do this, I'm going to use the asynchronous request (initiated by request method in the http library) to enable our "tour guide" and navigate the hazards by emitting an event (using the EventEmitter from the events library) after we've handled the response.

But...but...but...I heard that building anything synchronous in nodejs is wrong, and this sounds synchronous. Yes, in the application sense this basically turns asynchronous commands into a synchronous process, but only somewhat, as it's really more a synchronized process, not synchronous. After all, our tour guide doesn't need to know what else is going on downstream or what's going on under the water or what other rivers are like - they only need to know how to navigate this hazard on this river to get out safely on the other side.

Beside the fact that this is more a synchronized process than synchronous, we must admit that many things in the real world rely on being at least somewhat synchronous. For example, how can we know if we can update a profile unless we know the profile exists or is created or how can we know if we can ship a product unless we know we have it? We can't. Does that mean for all of these synchronous activities we cannot use nodejs? Poppycock. Yes, there is great power that comes with asynchronous processes - but that perception of power is because we assume the processes can be run in parallel. I could insert the old joke here about nine women having a baby in one month, but we all know the reality - we all recognize that there are some processes that cannot be run in parallel.

Another way to look at this is that sometimes, we lose focus and forget that while our single boat is navigating the river there are other boats and other rivers. If you're really concerned about performance, build your app in such a way that all boats on the river travel the river safely at the same time, even if each boat has to navigate from point to point.

Back to our hypothetical...a simplified portion of the code for this might look like...

function doNextStep(response) {
  if (response.status_code === 200) {
    switch (response.step) {
      case 1:
        doStep2();
        break;
      case 2:
        doStep3();
        break;
    }
  }
}
function doStep1() {
  // create the site
  destination.step = 1;
  sendRequest({host:myhost, path:'/bin/wcmcommand', port:4502}, destination);
}
function doStep2() {
  // update the site
  branch_data.step = 2;
  sendRequest({host:myhost, path:'/' + childsite, port:4502}, branch_data);
}
function doStep3() {
  // log the results
  logResults(branchData);
}
…
function sendRequest(options, data) {
  var req = http.request(options, function(res) {
    var body = '';
    res.on('data', function(chunk) {
        body += chunk;
      });
    res.on('end', function(err) {
        try {
          data.response = JSON.parse(body);
          emitter.emit('step_done', data);
        } catch(err) {
          console.log(err);
        }
      });
  });
  req.write(data);
}
…
emitter.on('step_done', doNextStep);

Of course you could use something like a timeout to build in human-like reaction times, and that might work for most of the situations, but the problem here is generally not on the end issuing the HTTP requests, it's with the end receiving the requests and processing them. We need something that watches for the response instead of waiting for a response, and this is why the callbacks on the asynchronous methods were built. Like a tour guide, the emitter calls out each time we've gotten a response and are ready to address another hazard as we've finished with the last.

In this example, we've created a river of events that our tour guide navigates, and in our hypothetical, synchronized process, it would be conceivable that we could set up a nice website in less than 200ms. The difference in speed of a (long) 200ms synchronized process compared to a human doing the same task makes me care less about whether or not something meets my anti-synchronous ideology.

Happy coding.