Tuesday, November 20, 2012

The Simplest Thing

I recently took part in a code dojo, using the scoring for bowling kata, where we were interested in intention revealing and expressive code. The rules set up were typical, TDD, rotating pairs, no verbal explanation to new pairs, no talking to your current pair. Same old rules, same old results, no surprises. It seems to me that if you want different code to come out of the activity you may need to frame the activity differently.

The same old results of TDD as practiced in the typical code dojo is not a good thing. Write a test and implement a solution to it as quickly as possible. Satisfies each test with the least number of characters possible and in the shortest time possible.

Short and quick lead to producing the easiest solution to come to mind, or perhaps simply the first solution that comes to mind, and basically the most primitive solution. By primitive I mean that the solution is implemented as directly as possible in the primitives of the language.

But it's TDD, the goal is to write the simplest solution. Not to be simple minded about your solution. This depends on defining simple of course, but in my experience that's a measure that should be applied to the result code, not the process that generates it. TDD is not about doing the simplest work that results in something that satisfies the test, it's about doing the hard work of delivering something which is simple and satisfies the test.

Simple becomes a set of principles, some with potential metric support, some more like rules of thumb. Simple implementations for me are those which can't be used wrongly. Simple implementations make explicit the important things, not merely the static things that are easy in whatever language you're using but also the processes and dynamic properties of the system even if you have to twist the language to make them explicit. Simple often means hiding and wrapping primitives so that they don't interact and that the myriad things they might do but shouldn't are cut off.

Simple very often means qualities that are not captured in tests, simple is the judgement of a professional saying that the passing test is just the beginning, diligent professional conduct requires me to attend to these other values.

What I found was that we could have written better code faster in one of two ways, (the code I wanted had a layer abstracting frames, and chaining them together as the game plays). If I'd thought long and hard I could have presented the tests in an order that made the abstraction an immediate and obvious benefit. Or, trusting our professional intuitions and being principled, we could have implemented the highest interface in terms that corresponded to the description of the problem, that DDD approach would have automatically led to a frames based abstraction because that's how the domain experts talk about the problem. Or I could have trusted my little rules of thumb: wrap your primitives, abstract away the language. Perhaps if I'd been able to talk to my partner we'd've been able to establish that quality foundation, but once people started rotating the pressure was on to just follow suit on poor beginnings.

I don't think I've ever written code where an implementation that jumped from the highest level interface straight into code using raw language primitives was the best result. When working with code like that, grief and suffering are typical.

TDD is a process that helps me write good professional code by freeing me up from tracking whether my code is functionally correct, instead I can give more energy to all the other professional values that need to be embodied in my code. But this is my heresy, TDD says nothing about the quality of the implementation, but a poor quality implementation is a terrible cost a company.

The problem with that common type of TDD kata was that it punishes us for attending to code quality and architectural concerns (remember those other things in XP: coding standards and the system metaphor, they acknowledged that TDD needs balance).

So how do you get good, expressive, professional, intention revealing code? No guarantees, but you need at least to be able to identify good from bad, you need the boldness to say no to the bad, you need to be allowed to implement better than the poorest conceivable solution.

TDD is a game, it rewards getting stuff done quickly and crudely, it also rewards getting things that function. People who want the process to deliver quality are sadly disappointed, people who like to tick things off and generally score points can use it to churn out piles of muck. People who feel that quality matters often end up ignoring it when forced to practice with rules mavens and game players. What's really need is counter balancing games that rewards the other qualities of good code, qualities that have great business value but may not be easily measured like maintainability or job satisfaction. Such games may look like leader boards showing the latest model code for the team to strive to excede, perhaps as an ongoing vote or outcome from retrospectives. It may be a checklist of points that need to be judged as part of code review. TDD doesn't stand alone.

TDD: Making Bricks or Building Houses

TDD has shown itself to be amazingly productive practice. With it's tight feedback cycle it's great for learning. Having very specific disciplines and methods mean some excellent tooling has been developed to support it. It's a natural fit for the small stories and task we favour for iterative development. It's strengths are also it's weakness, specifically, it drives an intense focus on the smallest units of development. And as time goes by it's becoming coming coupled ever more tightly to other practices that also focus down at small units of development.

That plays to a very human bias, substituting a hard question with an easier one. We puzzle of over a tricky question: are we making good software? TDD leads us to answer yes by substituting different questions: Have we written good units? Or, have we worked hard with sophisticated tools?

Good units do not make a good program any more than good bricks make a good wall or good musical instruments make a good orchestra.

The general tide in TDD is to keep looking down at small units. This is reinforced by mocking libraries and test frameworks and a culture that favours unit tests over test that aggregate units, (no one questions having a unit test but it is frowned on to have higher levels tests without unit tests, and slow running aggregate tests may well get axed in favour of a faster running test suite).

TDD teaching methods in some quarters have had a similar bias, discouraging conversation and discouraging looking further ahead than the next test. Further reinforcement comes from common patterns of pairing, ping-ponging for example. Likewise practices that focus on stories with little context, either chunked into iterations or striving to have just a few upcoming and potentially unrelated stories in a stream. Add to that developers who take pride in starting what has been presented as the next story without further consideration.

All these things make it easier to focus on the finest grain of problem solving. But practiced without balances they take focus away from the system as a whole being delivered, hiding synergies and commonalities and large scale structure. Indeed it can cause great frustration when someone practiced with the tools and techniques of TDD plows on into chaos dragging in their wake a partner who wanted more time to reflect or consider their work in a context broader than the current test or story.

In all this the mainstream of TDD risks being a "greedy reductionist". Greedy reductionists are a parody in science that says everything can be described in terms of the lowest level realties, basic particles etc, and that higher order abstractions are needless. Dennett gives an example that illustrates the problem: imagine a calculator that returns 3 when you press 1 + 1, everything else it does just fine; you can describe the operation of this calculator in terms of electrons and semiconductors and completely fail to explain it's most interesting and distinctive behaviour. TDD when focused on the smallest granularity risks the same mistake.

Focusing on small problems is called narrow framing and it can a powerful tool. In many situations the human mind instinctively, perhaps unavoidably, resorts to narrow framing. But narrowly framing problems can be deeply misleading about what makes a good decision. Our tendencies to over spend to guarantee a win or to avoid loss can be amplified in damaging ways if we don't take the big picture into account. Treating the parts of decomposed problems in isolation can exaggerate biases: TDD works against balancing our decisions in the context of the greater whole.

TDD helps us build good bricks, but it doesn't really say much about building good houses. Of course it's hard to build a good house from bad bricks, just there's a lot more that we fail to attend to if we only do TDD.

When the world was introduced to TDD it was in the context of system and practices that balanced those weaknesses. I personally have become very wary of programming with anyone people who don't actively practice techniques to balance the focus on the fine grain that is inherent in TDD.

There are many possible practices that can provide the balance; XP as a system has several built in:

The embedded customer and the system metaphor both help focus us on the whole system and it's values as a product, (sadly, they've often been the things left behind in emerging hybrid and customised development methodologies).

Code standards can be used to improve the focus on the whole and on code that works together as a coherent whole. One style of code standard identifies those areas of the code that embody the values at stake, (not just classes functions or files, but whole modules packages or pages that work as cohesive systems). Another style of coding standard uses guiding aphorisms, for example: business logic and events should be represented explicitly as objects; fan out should be limited; always wrap third party code; always wrap primitives and native library code. That last is can be argued but ties into the DDD principle of ubiquitous language: write layers and abstractions so that you are implementing in terms that reflect your users language. The quest for an over arching language rooted in the users expression of the problem domain, and having that language pervade the implementation can be a powerful tool for creating a coherent whole.

Perhaps one of TDDs greatest strengths is that it's easier to do than the counterbalancing practices and we naturally give energy to the things we do well or that are at least easy to do. Real growth, however, requires giving attention to what we do poorly. So the next time you go to do some deliberate practice or try to formalise your team's process, have the boldness to ask yourself what you need to do to improve or to cover for your weaknesses rather than just shining a light on your strengths.

Monday, December 19, 2011

Valid States Only, Please!

I've been thinking about why debugging can be hard and some higher level design heuristics that can help. One thing that really gets me down is invalid states. This problem has a slightly different face in functional programming versus object oriented programming. Sometimes it's not clear cut that a state is invalid, but rather we may have interim states that have no real business meaning.

Lets start with the classic list processing style of functional programming. It's pretty common to take a list of elements conforming to some design, they represent things in the domain in some given state, and the initial states are normally pretty sensible. Then the list is processed via a series of list processing stages, maps and filters, reductions, and so forth. The problem is that these primitives don't distinguish convenient intermediate steps that are just implementation details from stages that are really significant and meaningful states from the domain perspective, valid business states.

The solution is chunking together list processing primitives to make higher order list processing functions take lists of valid items and produce similarly valid results, hiding the intermediate workings, (valid here means representing a state that is meaningful in the business domain).

Failing to break the list at the important points leaves any future developer puzzling over the meaning and correctness of the intermediate lists, particularly if they want to change the process by perhaps adding additional intermediate states or shifting process so that subsequent states are different, or if they are wondering why an object arrives at a particularly stage of processing in a given state.

An analogous situation plagues object oriented programs. A method transforms an object, it does so via a series of intermediate steps. If the method makes those changes on the object directly then the object goes through a series of intermediate states. This can cause all sorts of confusion if you're trying to modify a transformation, it leaves the developer puzzling over what precise intermediate state the object might be in if they what to make change at some given point, effectively the developer has to run the whole program in their head to understand the context of any given sub part of the process. I'll just mention parallel programming, no need to say more I hope.

The solution here is to avoid changing parts of the object, instead, as much as possible calculate the new state in local variables and method parameters, then change the object in a single step. Two phases, one of data crouching and then another of state changing.

Too many months or maybe even years of my working life have been spent trying to understand the state of programs that allow wide access to intermediate states and don't distinguish intermediate states from significant stable states with real meaning in the business domain.

So, some rules of thumb:

  • Don't go tweaking state ad hoc, (getters and setters are for tools not programmers), state changes should correspond to meaningful business events.
  • Aim to hide, even from other methods, the intermediate states that come about as you implement a method.
  • Don't let raw list processing primitives become exposed in your business abstraction layer.
  • Chunk list processing primitives together into higher order operations which take you from one business state to another.
  • Know what constitutes a valid state, and don't let your objects ever be in invalid or intermediate states.

Thursday, December 15, 2011

Getting Stuff Done

I'm planning another blog to track my other interests: songwriting and music in general, and perhaps gigs and other social stuff. Anyway, I've been reading The Frustrated Songwriters Handbook and re-reading Writing Better Lyrics and getting some great ideas and general motivation from both. They are very different books in almost every imaginable way except one.

Both books deal with the problem of slow startup time for creative activity.

The central activity frustrated songwriters are encouraged to try is the 20 song game: set aside a day to write 20 songs and then show them to other songwriters who have done the same. Yes, 20, in one day - then do it again. I'm planning my first effort soon, maybe this weekend, if not then in the week after Christmas.

One of the very first exercise in Pat's book is called object writing set aside 10 minutes each day to write concretely about senses and experiences surrounding some particular object. Stop at 10 minutes. Even if you're bursting with ideas, stop at 10 minutes. After just a few days you start to learn to skip the pretence and centring and all your little rituals, all you can do is as he says dive deep for the good ideas fast.

Both these books are addressing problems of being blocked, of procrastination, of inefficient work practices, and such, that plague songwriters. It's particularly a problem for songwriter because it's normally something you need to fit around the rest of your life. Obviously as an amateur, but even for professional performing musicians distractions like preparing for the next gig, or taking students, and so forth, get in the way of writing music.

These songwriting exercises are not directly about producing songs, they're training techniques that help you hit the ground running when you write.

It seems to me that a very similar problem effects many professional programmers. There's a lot of code that I want to write, more than just what I must write in the office. I want to polish up portfolios pieces, I want to explore new ideas and learn new languages, I want to demonstrate all manner things, and I never seem to have the time, family and friends need my attention, household chores must be done, and so on and so forth.

So, as programmers, are there similar exercises that can make small chunks of time more effective for practicing what we do?

Here's a challenge, the equivalent of object writing, half-hour programming sessions. Make yourself a promise that for some reasonable period of time, (a week, a month, ...), you will work on a project for 30 minutes each day. At the end of 30 minutes, no matter where you're at, stop. If you're in the flow, stop, and let that remind you to get rolling even faster tomorrow. Maintain the commitment to 30 minutes every day for your chosen period, don't work longer today and then give up tomorrow. Outside of that 30 minutes you might need to fix up your tools and environment to make sure you can get the most out of that 30 minutes.

Here's another, the equivalent of the 20 song game. Write 6 programs in one day. They can be anything you like. It would be more interesting if they were different types of programs, maybe even different languages. But the should actual work as deployed programs. For example, write: a game, some data visualisation, a to-do list, a network monitor, a shopping cart, and a web photo light box, all in just one day. Can you wake with no idea what you're going to write, and just make a bunch of programs. Can you do it again next week? What bits of your environment will you need to polish and tune to achieve that 6 program target, what will you need to practice so that you don't waste time next time? Unless you think they are exceptionally cool, they need only be shown to other programmers who've taken the same challenge.

Good luck, here's hoping for a more creative and constructive future.

Wednesday, December 7, 2011

Red and Blue Architects

I've been thinking about what I expect from a software architect in the highly skilled egalitarian teams doing iterative development that I'm part of, where we're all able to make major changes in response to changing requirements and discoveries in a changing environment.

The traditional software architect's role has been to present a right way of doing something in advanced of stubbing your toes on a problem. That presumes we know enough about our product and it's environment in advanced to make those decisions. It also tends to under-rate the capabilities of other developers to solve problems. Let's call this lot red architects. That was actually true, and good architects in that style were valuable, when I worked in the hardware business: the experts really did have deep knowledge that wasn't available without experience of the product line and general business domain, and the operating constraints of the hardware really were pretty clear in advanced. I haven't worked in that world for many years.

Now I'm working in startup land, it's much faster and looser. The problems of bad architecture still plague us but architecture delivered by red architects just hasn't really worked. We need a different sort of architect, I'll call these new folk blue architects.

Blue architects don't tell me what to do. But they insist that a certain class of question gets asked and answered. Questions that couple the business needs to the hardest technical problems so that they can't be ignored and everyone sees the value in solving them. These are the guys who ask how we're going to handle cache regeneration when we go through database failover, or how we're going to separate our background report from our web front end when the data volume makes the background processes so costly they'll impact the servers responsiveness. They ask what we're going to do when we no longer have a window for downtime because we've become globally successful. Questions that steer us away from the superficial and primitive toward effective and satisfying solutions.

Blue architects guide the whole team to produce better implementations by asking penetrating questions.

Of course everyone with some technical nouse can ask blue architecture questions. But I expect the guy who claims to have architecture specialisation to be better at it, to ask deeper questions, and to frame those questions in ways that lodge in the team consciousness. An architect fails if they can't communicate their insight; I expect anyone claiming specialisation to actually deliver.

In the last few years of work the red architects haven't had much success. On the other hand I've seen individuals deliver blue architecture questions to great effect. Let's have more blue architects.

Saturday, November 19, 2011

Object orientation is more than a language primitive

Distinguishing object oriented languages from non-object oriented languages is like talking about free trade and regulated trade. All trade is regulated in both extreme and subtle ways; it's never free (think about slave trade, child labour, pollution control, safety, truth in advertising, and so forth). So called object orient languages provide primitives generally for just one object oriented abstraction; so for any given object oriented language there are many object oriented abstractions and features that it doesn't implement. All object oriented languages are just some person's or group's preferred abstraction (or merely easily implemented abstraction), not at all a universal truth.

I have a distrust of coupling my domain abstraction to language features. A lot of what a like about using functional style in imperative languages is that it helps me to recognise and explicitly represent some things that I might overlook if I just stayed in the mindset of the imperative language I'm working with. Object orientation abstraction is one thing regularly overlooked and left as a non choice. We generally fail to ask the question: what is the right way to use object orientation in this specific application?

It's important to me to recognise what different object oriented abstractions provide: different ways of expressing type relationships, different ways of sharing common code, different ways of protecting access to code and data, different ways to communicate between objects, different ways to get new objects, and so forth. With that understanding I want to build the most appropriate object oriented abstraction to get the benefits most relevant to my particular problem.

One tool to help reach this end is to sneer a little bit at the object oriented primitives in whatever language I'm using. By not letting myself lazily treat that set of primitives as automatically the right way to do object orientation I make myself ask valuable questions: what is the right way to partition code in this application, what is the right way of expressing type in application, what states or life cycle will my object pass through, and much more.

This has recently been brought home to me in discussions about programming Lisp. Programmers that I respect say they have no idea where to start writing a program in Lisp. I think they expect to have the language tell them to organise their code. But really in Lisp, beyond the beginners stage, you're job is to decide how to organise your code, how to represent object orientation if that's part of your application, or how to model data flow, or logical structure, how to represent what ever is important in your application.

In the JavaScript world there are a number of systems that offer support for building object hierarchies, providing abstractions that are generally more like Java's object oriented abstraction. This is a great thing if that object oriented model is the right abstraction for your client side programming. Sadly I think it's been done because people find it difficult to see object orientation as a choice of application specific abstraction rather than a constant of the tools they are using, and aim to make all their tools homogenous with regard to object oriented abstraction.

Oddly, the best object oriented code I've ever dealt with has been in C and Lisp, probably because the developers made choices about organising their code, and didn't treat the primitives of the language as the end of the story. Higher level abstractions of data flow, and of aggregating data and function, and object life cycle, were clearly chosen and expressed in code. Avoiding accidental design driven by language primitives.

Saying or acting as if there can be only one object oriented abstraction and one set of perfect primitives implementing it is frankly and obviously silly. There are many object orient abstractions out there being used productively every day. Good programmers should be able to choose the abstraction that's right for the application and implement it in the language that they are using for the application.

Tuesday, November 8, 2011

Dynamic dispatch is not enough to replace the Visitor Pattern

I've been bugged a little by how poorly understood the Visitor pattern is by people claiming it's not needed in their particular language because of some funky feature of the language. In the Clojure world dynamic dispatch or multi-methods are the supposed giant killer, (and I'm an Lisp fan of old, and Clojure is fast growing on me, this will be a rant about the fans who rail against an enemy and get it wrong).

The Visitor pattern has two parts, an Iterator and Dispatcher.

The mistake is that because Clojure has dynamic dispatch then it must be the dispatcher that gets replaced. In fact the dynamic dispatch really can help with the iterator, allowing different things in the structure to be treated differently. But as the example below shows it's not sufficient for some even more arbitrary dispatch types, not without adding extra information to parameters to facilitate the dispatch, which is an unnecessary added complexity; unnecessary added complexity being my basic definition of bad design.

My favourite example is a tree traversal engine, for something like an AST, with data like this: (def tree '(:a (:b :c :d) :e))

Now I want to do arbitrary things to this tree, for example, render it as JSON data:

{"fork": ":a",
 "left": {
   "fork": ":b",
   "left": ":c",
   "right": ":d"},
 "right": ":e"}

And also as an fully bracketed infix language: ((:c :b :d) :a :e)

There is absolutely nothing, but nothing, neither in the data nor the traversal that can be used to distinguish between rendering as JSON data or rendering for an infix language. This not a candidate for any sort of algorithmic dynamic dispatch based on any part of the data: the distinction is arbitrary based on functions chosen without regard to the data.

It's quite fair to note that there are two different types of nodes, branching and terminal, that need to be traversed in different ways, which can be usefully used for dynamic dispatch, but in that's dynamic dispatch in the traversal engine not the final functional/Visitor dispatch.

First define a generic traversal:

(declare ^:dynamic prefix 
         ^:dynamic infix 
         ^:dynamic postfix 
         ^:dynamic terminal)

(defmulti traverse #(if (and (seq? %) (= 3 (count %))) :branching :leaf))

(defmethod traverse :branching [[fork left right]] 
  (do
    (prefix fork)
    (traverse left)
    (infix fork)
    (traverse right)
    (postfix fork)))

(defmethod traverse :leaf [leaf]
  (terminal leaf))

Nice application of multi-methods and dynamic dispatch. Note that in the most recent versions of Clojure (1.3.0) dispatch defaults to a slightly faster common case of non-dynamic dispatch, so you have to declare the dynamic binding type. If you know you need dynamic dispatch you need to tinker with meta-data, and you can do that to other people's code from outside: (alter-meta! #'their-symbol assoc :dynamic true)

For ease sake I'm also going to provide a default set of behaviour, doing nothing. Notice the matching dynamic binding type specified as meta-data:

(defn ^:dynamic prefix   [fork] nil)
(defn ^:dynamic infix    [fork] nil)
(defn ^:dynamic postfix  [fork] nil)
(defn ^:dynamic terminal [leaf] nil)

Now to get it to do something real, rendering in a fully bracketed in infix style. Notice that binding doesn't need explicit dynamic meta-data, because dynamic is the heart of binding:

(binding [prefix   (fn [_] (print "("))
          infix    (fn [fork] (print " " fork " "))
          postfix  (fn [_] (print ")"))
          terminal (fn [leaf] (print leaf))]
  (traverse '(:a (:b :c :d) :e)))

How about printing the symbols space separated in reverse polish notation, as if for a stack language like PostScript or Forth. Notice how I only need to provide two of the four possible bindings thanks to the default handlers.

(binding [postfix  #(print % " ")
          terminal #(print % " ")]
  (traverse '(:a (:b :c :d) :e)))

The fully bracketed rendering illustrates one other very naive mistake with many peoples understanding of the Visitor pattern. Some nodes are the data to more than one function call during the traversal, the distinction being dependent on the traversal pattern and the nodes location in the structure. The final dispatch to the is not determined by the data of the node in isolation, the individual node data being passed to the function is insufficient to choose the function.

There are some useful applications of multi-methods in the final dispatch, for instance to distinguish between different types of terminals as in this example:

(defmulti funky-terminal #(if (number? %) :number (if (string? %) :string :other)))
(defmethod funky-terminal :other [terminal] (println "other: " terminal))
(defmethod funky-terminal :number [terminal] (println "number: " terminal))
(defmethod funky-terminal :string [terminal] (println "string: " terminal))
(binding [terminal funky-terminal]
  (traverse '(:a (:b 42 :x) "one")))

But that dispatch isn't the whole of the Visitor pattern, and frankly to hold it up as the Visitor pattern only demonstrates that you haven't grokked the complexity of dispatch from a complex data structure.

Consider another typical use of the Visitor pattern: traversal of graphs. Let's propose a traversal of a graph that is is guaranteed to reach all nodes and to traverse all edges while following edges through the graph. There are many graphs where this would require visiting some nodes and some edges multiple times. But there are many applications that would only want to give a special meaning to the first traversal of an edge or the first reaching of a node. There may even need to indicate the traversal has restarted discontinuously in a new part graph in those cases were there is not edge leading back from a node or where the graph is made of disconnected sub-graphs. Nothing in the node itself can tell you that this is the first or a subsequent reaching or traversal, much less that the traversal has had to become discontinuous with regard to following edges, (maybe you could infer that from reaching a node without passing an edge, maybe you could track nodes and edges, but why do all that work when a good visitor will tell you directly by calling a suitably named functions; if we're going to enjoy functional programming, lets not be ashamed to use functions in simple direct ways). All these different function calls represent different targets of dynamic binding in the model I've presented her. Of course you could use multi-methods to achieve dynamic dispatch if there are different types of nodes for example. But it would still be the Visitor pattern if there was only one type of node and no reason for dynamic dispatch.

Dynamic dispatch is a powerful feature in Clojure, but trying to justify it as a giant killer to the Visitor pattern shows a lack of understanding of Visitor pattern, and weakens the argument. Present the languages strengths, not your own weaknesses.

Friday, October 28, 2011

A few of my favourite programming heuristics

There are many aphorisms and heuristics that can helpfully guide programming. Here are four that mean quite concrete things to me, and that have been characteristic of easy to understand but practical and performant implementations during my career.

  1. Clear object life cycle
  2. Analysis and action are separate concerns
  3. Single direction of flow
  4. Hide your technologies

(1) and (2) are particular takes on the single responsibility principle, focusing on the system, at any point in time, be doing one thing. (3) has that in it as well, do one thing and then move on to the next; design so that you're not coming back with feedback and then branching in response to intermediate states. (4) is addresses the laziness that accompanies excitement about the shiny new stuff; it's new, it's great; what could possible be wrong with letting it touch everything, (languages and frameworks being the most dangerous shiny stuff because they are hard to contain).

Know the life cycles of your objects and make them clear. Consider an object system that involves building up a web of objects and then sending it request. It's really useful to draw a sharp line between the phase of building up the web of objects and then subsequently triggering operations. Keep them separate. Where possible try not mutate the web of objects so the system is less likely to wander into invalid arrangements. Keep mutation separate from queries about it's present state or asking it to perform it's work. Explicitly representing object life cycle stages and the transitions and causes of transitions between them is one of the most powerful types of self documenting coding style I've every encountered.

Getters and setters are the enemy of the well managed object life cycles. The classic (anti-)pattern of creating an object and then using setters to tweak it into a valid state means you created an object in an invalid state. You really don't want your objects to have stages in their life cycle when they are invalid; don't allow object tweaking: ban setters. Getters mean other objects can react to what ought to be this object's internal representation, making it hard to change that representation. In messaging terms freely accessible getters are one step away from public attributes, broadcast messages, with each one exponentially increasing the combination of state information available in your program; don't feed runaway complexity: hide state, ban getters.

Many objectives require analysing a situation and acting, but analysis and action are very different beasts and it's best to keep them separate; keep two clear phases that carry you to your goal. Separation of analysis and action is my reaction to the nasty anti-pattern of getting a little information, doing a little work, getting a little more information, branching off somewhere, doing a bit more work, getting some more information. That way of programming means you need to know the history of the processing that got you to the interesting point you're trying to understand. My preference is to take time to figure out the current state, then act on that knowledge. Planning complex mutations or the construction of complex new objects and object collections this way is excellent; like the builder pattern, gather all the information, then act on it; similarly the interpreter pattern, build up a set of actions, then execute those actions.

Single direction of flow, or pipelining, is more of the same really. Avoid going up and down the stack repeatedly, and avoid too much back and forth in complex dialogues. Tell the down stream process or function everything it needs, how to handle errors, where to send results; fire your message; and turn your back on it. It requires you to more explicitly represent different states and possible reactions to those states. Frankly it's easier to trace and debug, less of the nonsense of being in very different states from one line to the next, less needing to know the accumulated side effects of the lines of code in front of you. Further, if you start with a model like this it's much easier to scale via message passing systems, a real win if you're actually planning to be successful.

You use many complex technologies, keep each one well contained. Hiding your technologies is about keeping separate different kinds of complexity. Don't get trapped in a framework, don't let someone else's technology dictate how you organise and present your code. You have complex business logic, represent it clearly. It is almost never the case that your choice of technology is the same as your business logic, so don't let them blend together. Don't design objects that expose they way they are being persisted, (the fact that they are being persisted is clearly a requirement and should be represented, but not specific how).

We all have different experiences and abilities and they shaped our personal measures of quality and our personal rules for achieving quality. It's good to know the values you hold and what in your history led to you holding those values. These four principles deliver their primary value by making large bodies of code easier to understand and particularly easier to debug and extend. I've spent too many hours puzzling over code, particularly stressful hours debugging code or trying to feel safe putting in place what ought to be a simple fix. They work by making temporal interactions simpler, making it clear what has been done, what is known, and what remains to be done. They're shaped by both success and failure at working on some very large bodies of code, and working on the same bodies of code for long periods as I tend to keep my jobs for several years.

Friday, August 12, 2011

Low Level Leadership

By dint of having more years of experience than many of my colleagues I get to be a leader, not as a formal position, but as an inevitable fact of humans in groups. I've been thinking about this for a while and I wonder about what a leader from within should be aware of in order to support those around about. These questions capture some of the things that I try to notice about my co-workers.


  • What do they fear most, or what makes them insecure in their work?
  • What would this person like to do better?
  • Of the things that need to be done what do they most enjoy?
  • What do they find most disappointing?
  • Should I or could I do what I've recommended that they do?
  • What grabs the attention of this person and what do they miss or ignore?
  • What factors or environment surround their best work, (collaborators, location, tools, time)?
  • What do they think is their best work?
  • What things most regularly slow this person down?

There's another angle to good leadership and that's to model good practices and behaviour. So I think asking myself these questions would be a good way to spend some time this weekend.

Thinking about some things my personal trainer in the gym has been asking I probably need to help him answer similar questions about me as my leader in that context.

Saturday, May 14, 2011

The Right of Veto

I'm trying to get my head around the question, "Why do good developers write bad code, and what can we do to fix it?" It's provoking dozens of ideas. This one came to me yesterday:

All developers should have the right of veto on any code in development. Developers should be taking the time to look at the rest of the team's work, to ask the questions, "Do I understand this?", "Could I maintain this?", and "Do I know a better way to do this?". If there's a problem in understanding, or a better way to do it, then exercise the veto.

Exercising the veto means that code doesn't get committed, or should be reverted, or must be improved.

Exercising the veto means you must now get involved in making it better.

Pairs of developers can't pass their own code, they need to get someone else to consider their code and possibly veto it.

No one is so senior as to be above the veto.

If there is disagreement between the code author(s) and the developer exercising the veto then don't argue, bring in more of the team. Perhaps the resolution will be to share knowledge better rather than the veto. But ideally you want to write code that any other team member will readily accept.


As I was thinking of this and talking about with my pair I realised we were writing code that I felt I would want to veto, so the next idea is that every developer or pair announces the state of their code. I plan to stick a big green or red button on top of my monitor; the red button means I veto my own work, and green means I think this is okay.

Self veto, (the red button), is a visible an invitation to help, work with me, pair with me, take me aside and talk about design or the domain or our existing code base. Open to veto, (the green button), is a request, please look at my code, judge it, share in the ownership of it, take responsibility for the whole teams output.

One of the threads that led to this idea came from circumstances that have led to us pairing less. My response was to ask what else we could do to improve code quality: how do we conduct worthwhile code reviews.

I'm heading toward the idea that even pair programmed code could benefit for code review, but I want to make it as light weight as possible. And the real point of code review, the thing that makes it meaningful, is the possibility of rejecting the work. Otherwise it's no more than a poor tool for sharing knowledge.

That realisation that code review is most valuable if it can be acted on tied in with a quote I read, "If you can’t say “No” at work, then your “Yes” is meaningless." The simplest "no" in code review is the veto on commit.


When working in sprints or iterations there's often a step where the developers commit to doing a set of tasks. But I've never really been in a position where I thought I could say no. The best has normally been to say that it will take to long or to negotiate details, but at the end of the day we assume that the right things are being asked for and we try to do them.

It's not an ideal commitment if I can't say no. Can my commitment be real consent if I can't say no? Certainly there have been times when I thought the requirements were unwise, and I've been demoralised doing what I thought to be the wrong thing. Could we create a team, or a work flow, or a process, that really empowered developers to say, "No, we cannot commit to doing that", and be honoured for it?

Wednesday, May 11, 2011

XP is more than Pair Programming and TDD

There are 12 practices in XP and 5 values. The values are open to interpretation but the practices are pretty clear. Many people magpie practices and while the practices are often useful in isolation there is a synergy between them which can be lost.

Coding Standards is one of the practices. To me it encompasses more than the simple things captured by findbugs, checkstyle, and source code analysis tools. It goes beyond naming standards, line lengths, function and class size, and even sophisticated measures like fan out, cylcometric complexity, and whatever other panaceas.

Coding Standards should bring us toward principles like: any developer should be able to understand the code others have written; our implementations should correspond to the way we talk about our system; a developer should be able to figure out requirements from the written code; it should be easy to represent or articulate what we have implemented so that we can meaningfully reason about it and reliably build on it.

It's frequently not so. Frequently TDD and Pair Programming produce code that only makes sense to the developers writing it while they are actively working on it, sometimes developers don't understand what they're doing even as they satisfy tests. Tests frequently capture behaviour that is incidental to requirements. Developers hunker down and satisfy requirements expressed in tests but don't look up beyond that narrow focus.

Failing to attend to Coding Standards is failing to do XP, or at least failing at the whole cloth, and undoubted exposing development to greater risk of producing poor quality results.

There's nothing wrong with building up a personal or local process by picking elements from different development systems, but sometimes it feels like those personal processes are about making life easier by rejecting or neglecting things that require active practice to get substantial quality.

Coding Standards, don't leave home without them.

Monday, May 9, 2011

Implicit and Explicit

Each step of language design gives us tools that make more behaviour implicit, except some things that languages make implicit are really requirements better made explicit. So another thread in language design is to make it possible to be explicit about the things that matter.

The three basics of structured programming are sequence, selection, and repetition. I'm going to use them to discuss how languages make things implicit or explicit.

Selection

The basic tool of selection is the if statement which connects different states to different code paths:
if (condition) then

fnA()
else
fnB()

The condition inspects available state and in some states selections fnA and in others fnB. This is an implicit relationship. It is a coincidence of placing these statements in this way. If we had to make the choice again in some other piece of code we could easily get it wrong, we could construct the condition badly, or forget we had to make a choice at all and just call one of the functions, or perhaps not even have access to the state required to make the decision.

The use of polymorphism in typical OO languages makes the connection between state and function explicit.
obj.fn()

The relevant state is represented in obj and how it was instantiated. There are in principle different types of obj, but for each one there is a correct version of fn and that is the one that will be called. It is explicit in the fact of obj being of a particular type that a particular fn will be called.

Repetition

The basic tool of repetition is the loop:
while (condition) do

fn()
done

But we do many different things with these loops, we might for example iterate over a collection of numbers and sum them. Or we might make a copy of that collection while translating each value in some way. Many languages provide constructs that make the meanings of such loops explicit while taking away all the accidental machinery of loop conditions and so forth.
collection.reduce(0, lambda a x: a + x)

or
collection.map(lambda x: fn(x))

Given constructs like these there's no possibility of missing elements because of the condition being wrong, and by drawing on the well established vocabulary of list processing the meaning is made clear.

Sequence

This is the tricky one, it's an implicit behaviour to which we are so accustomed that we rarely notice we've relied on it. Many bugs happen through accidental dependency on implicit sequence. Consider:
 def fn()

stat_a
stat_b
end

stat_1
fn()
stat_2

We all know that stat_2 is the statement that will be executed after stat_b. First, what else could it be?

In a parallel environment if could be that statements at the same level in a block are executed in parallel, perhaps interleaved, but with no guaranteed order. Simple as that, and very important given the rise of multicore devices.

In such an environment we would need to make sequence explicit. Consider a construct that forces consecutive execution in a parallel environment:
stat_a ::: stat_b

So stat_a will be followed by stat_b. You might also find things like this in functional programmings in the form of monads.

Another form of explicit sequence is the continuation, a mechanism for representing where the next line to be execute is to be found, and a matching mechanism to actual go there. These mechanisms often look like a type of function call, but one for which there is no notion of return, which results in no need for a stack. I admit continuations are one of the weirder twists in computing languages, it can be hard to know when to use them, but I've found they become useful in patterns where you manage the calls, perhaps for deferred execution, or to trap exceptions.

There are other interesting constructs that can help express important, required, sequential operation. With resource management that execute a block of code and then guarantee that after the block something particular will be called. For example the automatic try with construct in Java 7, or perhaps yielding to a block in Ruby.

Don't look at me like that...

It might seem that all the simple forms that I'm calling implicit are spelling out the steps more clearly and must surely be explicit. The trick is to think in terms of required behaviours: if you have a required behaviour how easy would it be get it wrong with the simple forms? A statement shifted to the wrong place, a sign wrong in a condition, and generally nothing in the simple form to tell you what the right form would have been, what was the intention of the construct. The meaning is implicit not explicit.

Conversely, if you use the explicit forms when sequence, repetition, or selection are truly requirements and you will be announcing your intentions very clearly. You'll be making it easier for libraries and languages to do clever optimisations to make use of parallel processing or for safe and efficient use of memory and other resources.

Tools and techniques that work against us

I'm a keen fan of both agile methods and modern software tools but I've become more and more aware of how they can also work against us.

In the Java world we have Eclipse, a fantastically productive IDE, but it has some features that most developers rely on that work powerfully and inexorably against good design. First, it defaults to hiding your import statements. Second, the auto-complete looks outside the currently available scope for possible matches. Third, it makes adding import statements nearly automatic.

The net result of Eclipse's help with importing is to circumvent Java's quite deep and powerful package structure. You are exposed at all times to all the classes in your application. And people love quickly importing any convenient class, programs lose structure, developers depend on knowing class names rather than judging what they should use from things that have been considered and made available deliberately. Everything is connected to everything forming a black hole of complex interconnections that sucks in vast amounts of developer effort.

Slightly less obvious is ease of adding public methods, (indeed Eclipse will offer to make methods public for you), and by default decorating them doc comments, which means that momentary conveniences expand interfaces, creating more ways to couple different parts of your system together, and accreting functions in interfaces with no sense of those things being right for the API. That momentary convenience creates something that looks deep and intentional but isn't. Of course it's made worse by those lazy agile practitioners who jump at a functional solution now and either neglect or are ignorant of the whole application quality and direction.

Between the power of tools to work to give easy access to anything in the system, and the deliberate small scale attention focusing of some Agile practices, with have a powerful combniation which we can use to quickly dig deep holes and traps.

Sunday, May 8, 2011

Continuation Passing Style in JavaScript

One of the frequently unquestioned assumptions in modern programming languages is "what comes next". We just assume it's the line after the current line, or the first line of the function just called, or the line that follows where the function was called from.

Continuations are objects that say go to this point in the code, and there is no return, where next becomes explicit. They are not goto statements, they are objects that can be passed around and represented in structures. A common example is using continuations to model exceptions but with no requirement on a continuation to be somewhere "above" in the stack.

A continuation syntactically looks a bit like a function but with no going back there is no need for a stack frame. Function call without a new stack frame is a key concept in tail call recursion optimisation. Languages that support one often support the other.

But approximating the style can be useful even in a language which has no explicit support for them, for example JavaScript.

A JavaScript Recursive Descent Parser in Continuation Passing Style

To illustrate I've transformed this straight forward, hand crafted JavaScript, recursive descent expression parser to continuation passing style and along the way also into tail call recursive form. Both versions translate a token stream into tree, (useful for fully bracketed rendering, or evaluation with correct operator precedence).

The transformation means I can not rely on the return values of functions. Instead I can only go forward. I go forward by calling another function. So, the original version calls xpn, expects tokens to be mutated, passes the return value to term_tail along with the mutated tokens, and then returns that result.
 function term(tokens) {
return term_tail(xpn(tokens), tokens);
}

But translated it simply says to continue with xpn, then term_tail, and finally whatever c might be.
  function term(a, tokens, c) {
continue_with([null, tokens], continuation(xpn, term_tail, c));
}

And I start the whole process by saying at the end continue with done which prints the results. What will end up happening is the recursive tree will become a pipeline.
  continue_with([null, [10, '+', 11, '*', 12]], continuation(expr, done))

You can see I have a special function to prepare my continuation, and another to call them, we're deep in higher order programming territory here... But they're not deep magic, mostly they handle packing and unpacking arguments.

Because there is no coming back from the continuation all results have to accumulate forward in the parameters, and the continuation calls have to be the last thing done in the function, that bit is the translation to Tail Call Recursive form. Where the return of one function would have been used by another function call, that second function becomes a parameter to the first:
  fn2( fn1( 123 ))     // simple style
fn1(123, fn2) // continuation style

The last bit of cleverness is that you can insert other continuations in front of the input continuations. In effect, this algorithm works by taking the first function and the last function and squeezing all the other function calls in between.
  function xpn(a, tokens, c) {
continue_with([null, tokens], continuation(factor, xpn_tail, c));
}

So xpn is supposed to continue with c but has squeezed factor and xpn_tail in front.

Why Bother in General?

In a good functional programming language this style works without building up anything on the stack. And it's working without actively consuming tokens, allowing immutable data structures for the tokens. That means it can be very memory efficient, and also produces algorithms that are more suitable for parallelisation.

I also have a philosophical preference for being explicit about sequence when sequence is a real requirement, and practicing this style gives me another tool for being explicit about things that are often just assumed, (more on this soon).

As a general bonus tail call recursive functions can be trivially translated into loops which might be important if function call overhead is a problem for a given algorithm.

Why Bother in JavaScript?

The naive version in JavaScript, just modelling the continuation call with a simple function call, is probably a bad idea. It ends up taking all the branches of recursion and glueing them together in one long pipeline; a recipe for running out of stack space. But instead of simple function calls I've used setTimeout, so this potentially time consuming operation is now neatly broken up allowing the UI to remain responsive.

Because I'm managing the continuations calls explicitly, so I could trap all exceptions if that was important, just by changing the functions I use to build and call the continuations. Trapping exceptions this way is a great trick for concurrent programming because you can pass the exception handling continuations from thread to thread even though you're dealing with separate stacks, (you'd have to provide exception handlers as callbacks or additional continuations, a pattern I personally prefer anyway).

And as mentioned, I could translate all this to a loop, and do away with the function call overhead.

Saturday, December 25, 2010

Software development is all design work from a Lean point of view

A criticism I've heard made about some major voices in the Lean Software Development community is that they emphasis the reduction of waste. Waste is most commonly viewed as rework, or in the throwing away of spoiled goods, although there is also waste in opportunities from not getting quickly to market. A naive view of software engineering would treat refactoring and abandoning prototypes as waste, and consequently avoid practices that lead to those things.

However, everything that I've read so far emphasises some balancing principles in Lean that are part of making it work well in software development. The key is recognising, as traditional Lean does, that design and production are different domains and that different values need to be prioritised. Specifically, rather than a slavish and naive view of waste, there should be significant attention given to making decisions at the last responsible moment and having real options. 

Real options can be seen in the response to problems: If I come to you with a piece of work that isn't acceptable you could just say "no", or "that's wrong", or some other simple negation of the work. But much better from a Lean perspective is to give me more parameters for right, which in software might mean pointing out further requirements, or pointing out tests I need to be able to pass. There is no need at all for the negative part of the response. In fact, even if you follow up with all the good points, the simple fact that you've started with a criticism of my existing work will lead me to defend myself, lead me to try and validate myself by rejecting your ideas. If it actually matters to you that I respond to your advice then don't put me on the defensive, instead empower me with more knowledge and real options.

Software development benefits from making use of the last responsible moment to decide, rather than either locking in decisions early or avoiding making decisions until there are no options left. The later the decision, the less time between deciding and delivering, the more likely that decision is to be based on all the relevant facts and to be usefully out there before new facts turn up.

One of the big insights of Lean, and particularly important for Lean Software Development, is that real option are the things you need to keep available in order to make any decisions.

So I've been thinking about this, how can I maintain real options during code development so that I can defer choices until I have more information and choice closer to deployment?

Lean itself offers some good answers, like make it easy to revise decisions (refactoring) and having side-by-side prototypes. Refactoring shouldn't be seen as fixing mistakes or just tidying up at the end, it's the exploration and refinement of better designs and implementations after the requirement is understood, we do it to improve the product because without those improvements the product isn't acceptable to the users or becomes to expensive to maintain over time. Lots has been said about refactoring so I won't go into it more than that.

When I say prototypes in software I don't mean short lived experiments to be thrown away and subsequently reimplemented with better form, I prefer to call those things spikes and they are valuable learning exercises. What I really mean and want to talk about are side-by-side prototypes as real development options carried on in parallel, prototypes as the output of the design phase which suggests all developer work is with prototypes.

Sounds expensive, developing two or more real versions in parallel. More pragmatically you can try to keep real options about decisions in your domain that are easy to revise, like working up many user interface wireframes while keeping a minimal skeleton UI in place in code, or working up several different object designs on paper as Eric Evans encourages. But at some point you should try this with actual code, I've found it enlightening. 

In order to have parallel implementations in code, without the pain of branches and so forth, you need to be able to switch between both versions. That might be as big as having two versions of a page or dialog, or it might mean having a system of options and switches that allow you use one or the other. 

I'm finding there's an obvious benefit and a very important less obvious benefit of actually trying this. Obviously, I get to look at alternatives side by side. When I make a serious attempt at both alternatives the real requirements start to stand out from implementation details, so in the end the losing version informs the winner. But beyond that the code around the alternatives ends up improved. In order to plug in alternatives, and switch between them quickly, there needs to be a well defined and coherent place for them, a sensible provision of services to them, and a clear notion of what role they play. So, as well as improving the quality of the specific code being developed, side-by-side prototypes become a force for quality at higher levels and larger scales in the software.

Right now to me that would be a win across the board. Clearer design would make it easier to integrate new staff as we grow, real options would give them the chance to prove themselves rather than slogging down what at times seems the only possible route.

Wednesday, December 22, 2010

JavaScript: Meta-programming Example

Currently in vogue is to have code writing code, this is often called meta-programming. There are many different techniques, and many objectives when applying the technique, and it's often very powerful.

Here is a scrap from a little expression engine I wrote in JavaScript that uses a map of regular expressions to represent the various tokens.

    var lexeme_patterns = {
"integer": /^\s*(\d+)/,
"plusminus": /^\s*([\+-])/,
"multdiv": /^\s*([\*\/])/,
"open": /^\s*(\()/,
"close": /^\s*(\))/
};


There are a number of functions that I want that correspond to each of these tokens. For example I want a function that that I can ask if the next token is of a particularly type: lexer.is_plusminus() or lexer.is_integer().

Rather than manually construct these functions, I generate them using new Function(..) during the construction of my lexer. That way when I add a new token I get the functions I want automatically:

    function extend_for_token_set(proto, tokens) {
var is_token = "return this.current_token['type'] === '!!!';";
tokens.forEach(function(token) {
proto["is_"+token] = new Function(is_token.replace("!!!", token));
});
}(lexer, _.keys(lexeme_patterns);


Is it any use? Is it safe? Well, I find it useful to eliminate manual steps as I evolved my expression engine, and if I wanted to change behaviour there was a minimum of code to change, while still getting all the function I wanted for an easy to read style in my parser. But! it is tricky to figure out what code like this is doing if you don't already know what it's supposed to be doing, and that is a big negative.

So, good or bad, I don't want to say, but I do think it's worth knowing how to do it; it's a powerful tool in the modern programmer's tool kit.

Wednesday, December 1, 2010

XP System Metaphor - Redux

Talking to people, it may be that I'm taking the idea of the metaphor in XP too seriously. That it was an underdeveloped idea about architecture that amounted to little more than: We should be able to easily talk about your over-all system. We've learnt a lot more about the importance of language in design through Domain Driven Design. We should be aware of the history of it all and the lessons we're learning about how to do these things better. We should pay attention to things that are hard to talk about and bring them into our discussions.

That said, my experience is that a big system is much more approachable if you have a solid abstraction of it as a piece of software in terms that can easily discussed. Bonus points if you can keep clear the boundry between the concerns of making an application and the concerns of the particular business domain. And a special secret prize if you can use concrete language to talk about the bits of your system that don't correspond to anything in the outside world. I've had success with a metaphor for the system that is a distinct contrast to the business domain.

Tuesday, November 30, 2010

XP System Metaphor

The metaphor seems to be the most puzzling of the 12 XP practices. I'd been resistant to it because it seemed to obscure matters.

I'm quite taken with DDD and for me the heart of DDD is a well protected model that captures domain knowledge using the language of the experts. It seemed that the XP metaphor contradicted DDD's notion of ubiquitous language because the examples I'd been shown tried to replace the language of the domain with the language from the metaphor: the suppliers in our system are cows on a farm... no, not really.

Now I'm beginning to think DDD complimented by a system metaphor may be a brilliant solution. But it will depend on using the metaphor for the right part of the program.

I think the metaphor should be sought out for the parts of the system that are not clearly business domain, and further I think the metaphor should be strong and quite different to the domain. The metaphor represents the system architecture in all the ways that it is independent of the domain so I've started calling it the system metaphor.

With a good system metaphor the boundary between domain and system becomes clear: when the language of the metaphor and the language of the business mix there is an obvious clash of images. With a good system metaphor the architecture is made explicit in code: architecture is no longer an accidental fact of using certain libraries, accessing certain services, and calling one thing after another. With a good system metaphor explicitly representing the architecture and properly tested like any other code the system architecture should made easier to change and evolve independently of the domain.

Ron Jeffries gives an example of an agent-based information retrieval system using a beehive metaphor. The domain of information retrieval would use the language of information retrieval experts and those things which were part of the agent system and not part of the information retrieval would be coded using metaphor language; you know you're dealing with the system and not the expert's domain because you can see beehive words. Of course if the agents and the information retrieval were intimately linked in the domain then I'd avoid the metaphor and stick with domain language, but then I'd be looking for a metaphor to capture the system architecture around them.

Before I ever seriously looked at XP I had some good experiences with a distinct system metaphor. I was building a modular system for network printing, allowing specialist modules be plugged into drivers or moved to the printers themselves. What I built was pipe system, and in discussions and code I emphasised the metaphor of plumbing.

The plumbing metaphor allowed our specialist users to stay focused on their domains (colour processing, optimising streams of graphics primitives, font handling, etc) and to see their system needs differently to their domain problems. Thanks to the clear over arching structure that came from the metaphor, the program was very accessible to new developers. It was easy to explain, we'd draw pictures of pipes in sections that sparked imaginative thinking, (what happens if I change the order of pipe sections, or loop back on myself...), and it facilitated testing, (tapping the flow from a specific driver and pouring it back in elsewhere for generating test data, which later evolved into a network transmission system). We were able to effectively refactor the architecture without much fuss (changing the transfer mechanism between pipeline stages to improve error handling without touching the domain modules because the abstraction had kept them nicely separated).

But how do you find a system metaphor? You go looking for it. You also create the space for it. If the system metaphor is supposed to help you keep the system and domain separated then strive to keep them separate, try to talk about them with different language. If it's to makes sure that the architecture is something appropriately tested and structured then strive for those ends. And look for language which keeps the domain and system separate.

Be conscious of what goes wrong when system and domain get tangled, business logic in your controllers, changes to the user interface that break business functions, things that mostly happen because we don't distinguish between the domain and the system facilities we're using to deliver a program around that domain.

In the end it's a creative step, so get creative, try drawing the system without any boxes and arrows and without any words, think poetically, explain it to someone you respect who isn't a programmer.

And if you've had any success in this area, or if any of this leads you to write better programs please get in touch, I'd love to hear your stories.

Thursday, November 25, 2010

JavaScript: Curry Does The Work

curry is a common higher order function which isn't part of JavaScript on most browsers. However Prototype, Underscore, and many other libraries do offer it. The most up to date JavaScript implementations offer the bind function which can be used for the same job. I'm not going to show you how to implement curry nor attempt a generalised lesson, this is about how curry helps me with a refactoring.

In my previous post there was some really clear redundancy in the two wrapper maker functions. Lets make try to make that go away:
 
function wrapper_maker(before, after) {
return function(wrapped) {
before();
wrapped.apply(null, arguments);
after();
}
}

Which is great for the context push wrapper:

var contextualised = wrapper_maker(context_push, context_pop);

But how are we going to get those arguments for logging functions in place? The function curry does the work, (I'll use the Prototype version, it's pretty typical).

Basically curry makes a new function that behaves as if the arguments given have been pushed in front of other arguments so the following two bits of code do the same thing:

logger('abc');

var loggerX = logger.curry('abc');
loggerX();

This can be powerful when you want to pass a function that is generalised with parameters to another function that expects a function without parameters, which is what we need:

var logged = wrapper_maker(logger.curry('start'), logger.curry('done'));

And just like that curry gives a simple solution. It works when we want to add parameters to callbacks but don't control the code that will call our callback.

Refactoring to make our wrappers with wrapper_maker means we can add functionality to all of them without the duplication we would have needed in the previous post. Here's a refactoring that makes wrappers that pass back the wrapped function's return value:
 
function wrapper_maker(before, after) {
return function(wrapped) {
var result;
before();
result = wrapped.apply(null, arguments);
after();
return result;
}
}

Tuesday, November 23, 2010

JavaScript : Functions All The Way Down

Sometimes doing work with functions makes for more lines of code, but it also makes patterns of function calls reusable.

Here's an example that sandwiches our function call between some logging and some context management:

logger('start');
context_push();
do_stuff(data);
context_pop();
logger('done');

This kind of code gets dangerous if you have to repeat it. The problem is that we really want precisely that order of the surrounding functions. It's easy to transpose lines, or to update the pattern in some places but not others.

Functions come to our aid allowing us to write something like this:

logged(contextualised(do_stuff))(data);

That will guarantee the sandwiching, the first call will be matched by the last call, the next inner pair will match, and so forth. But actually writing those functions will add some code.

function logged(fn) {
return function() {
logger('start');
fn.apply(null, arguments);
logger('done');
}
}

function contextualised(fn) {
return function() {
context_push();
fn.apply(null, arguments);
context_pop();
}
}

The real value here is reuse. We can easily wrap any function now. And we don't have to call our function immediately, we can take the wrapped result and pass it as a parameter where ever it might be useful.

We can even make a general purpose function that create tools like logged and contextualised. I'll show an example of that when I talk about currying functions.