Tuesday, February 25, 2014

Modest Expectations

From Chapter 2: The Mythical Man-Month

Excerpt
This then is the demythologizing of the man-month. The number of months of a project depends upon its sequential constraints. The maximum number of men depends upon the number of independent subtasks. From these two quantities one can derive schedules using fewer men and more months. One cannot, however, get workable schedules using more men and fewer months. More software projects have gone awry for lack of calendar time than for all other causes combined. (p. 26.)
With those words, Brooks wraps up the pseudonymous chapter of his classic text. If there is any law for managing a software project, it's Brooks' Law: Adding manpower to a late software project makes it later. The mechanism is well understood and the basic concept broadly accepted—for decades. And yet, in program after program and project after project, the law is re-proven as if it was a traditional student experiment in a first year college lab.

I'm reminded of Santayana's often cited admonition:
Progress, far from consisting in change, depends on retentiveness...when experience is not retained, as among savages, infancy is perpetual. Those who cannot remember the past are condemned to repeat it. "1
If you think about it, not all history is repeated. We're particular good at repeating the mistakes. Why not the successes? How does that happen? Specifically, how does it happen in NASA?

The the other night, I was having dinner with my good friend 'TJ'. He was reminding me just how cool it is to work for NASA. People are living in space. A rover is driving around on Mars. We are getting images of exoplanets. His enthusiasm was genuine and infectious.

International Space Station (May 2011)
He's right of course. These things are cool, but they seemed tepid compared to what NASA did during the Apollo era. Is driving a Martian rover 100 yards comparable to walking on the moon? Does the first space walk compare vaporizing a Martian rock? Does the return of the Lunar Excursion Module (LEM)and re-docking of the LEM after a moon walk compare to the maiden docking of Dragon Module?

Back then NASA engineering was breaking new ground. No longer. We build spacecraft the same way we did 30 years ago. The new systems are still fragile and expensive to operate. And what about the ongoing existential crisis regarding the purpose and value of the spacestation. Does it make sense to spend more than $3B per year to keep astronauts on board to see if vegetables will grown in space? Is this headed anywhere? What's happened to the agency vision?

Later that week, I met my good friend, 'KN' for lunch. We had worked together on many tasks and shared disappointments from working on the wrong side of the political equation. That day he was surprisingly optimistic. He seemed to think that, at long last, a meaningful change was at hand. His project has committed to use model-based engineering techniques.

Over the years we'd seen similar decisions come and go. 'KN' is an old pro, no Pollyanna. Perhaps this time it would be different. Perhaps this was no mere lip service. Perhaps when the coding starts, these new techniques will actually be used. This was a glimmer of hope. But, as I thought back on the effort needed to obtain management consent for this small innovation, I wondered why significant engineering advancements for a NASA were merely routine decisions elsewhere.

480914main ELHVrailartistconcept
What was imagined in the 1970s
Back in the 70's, expectations were different; great achievements were expected. NASA was confident and flush with success. Routine and cost effective access to space was surely in reach. NASA built the shuttle. In order to travel past the Moon and beyond, we needed a jumping off place. NASA started work on a space station. These were bold objectives. Humans were about to push into the frontiers of space and NASA was in the lead.

That's when the setbacks started. The shuttle was delivered 2 years late. There was a tragic accident. The space station project became a muddle and nearly abandoned. NASA had overreached. Doubt crept into the culture. By the late '80s the Agency's growing conservatism slowed the pace of innovation. The engineering values shifted to predictability, risk reduction, and cost savings. Along the way, something changed—expectations were no longer great.

In the last decade the government tried to reenergize NASA with the Constellation Program (Cx). The goal of Cx was to return to the Moon and enable travel to Mars. From the outset, the program was cursed by an unrealistic funding profile, a fanciful schedule, and the rigid bureaucracy that sprouted since the glory years. The result: Cx floundered for 5 years before being mercifully cancelled. However, despite these crippling constraints, Cx did leave a legacy in the Orion Multipupose Crew Vehicle , the Space Launch System and the Commercial Orbital Transportation Services. These programs produced sensible vehicles that are basically Apollo-era recreations with modest improvements. Modest; sensible; but not especially inspiring.

Orion-spacecraft
The new Orion Multi-Purpose Crew Vehicle
Aside from occasionally bouncing or dangling rovers to the surface of Mars, the Agency's technical feats do little to animate the public's imagination. Now-a-days, it is the high-tech rainmakers, and not NASA, that feed the public's space appetite. Commercial space tourism companies tout the glories of a sub-orbital dip into space. Space cults recruit romantics who are clamoring to board a one-way trip to Mars. There are raft of grandiose projects, like space elevators and asteroid trappers, whose proponents proclaim their practicality with only the slightest consideration of the potential for complication.

It is this drift away from a culture of inspired, innovative engineering that is telling. Both 'TJ' and 'KN' are realists. They are professional. They are not complacent. They have internalized what's possible in the engineering culture of NASA.

Internalization of culture is a key to NASA's repetition of merely modest and just sensible engineering achievements. I saw a transformation in nearly every new hire. Most come on-board agonizingly eager, talented, and bubbling with ideas. They know little of the realities of engineering in the space business. In a few years most absorb the conventional wisdoms that engineering judgment rests on current practice and innovation lies on the borders of the status quo. They learn that engineering between the lines is the mark of professionalism. The result is an inward-looking, myopic culture busy with immediate concerns and reluctant to tamper engineering custom—even when that custom is to add programmers to address schedule problems late in the project.

Spider Space Station Concept - GPN-2003-00095
"Spider" concept for spaces station
 built from Shuttle hardware
There will be some fundamental assumptions which adherents of all the variant systems within the epoch unconsciously presuppose. Such assumptions appear so obvious that people do not know what the are assuming because no other way of putting things has ever occurred to them.2Whitehead
We are rewarded for internalizing the conventional assumptions and practicing the conventional methods. The problems occur when the circumstances,that gave rise to a convention,are no longer relevant. Knowing history is not enough to avoid the errors of the past.

I'd like to think I'm an optimist—perhaps one whose hope has been attenuated by disappointment. I don't believe we must repeat history or be mastered by convention. We can question fundamental assumptions. We can develop new approaches that will enable us to build the systems that were imagined in the 70's. If only a few remain tough minded and work to think past the assumed truths, 40 years from now the next generation will look back at the Mythical Man-Month and say: "It's a good thing we don't do THAT anymore."

With that, I'll move on to the next chapter: "Surgical Team."


1. Santayana, G., "Reason in Common Sense," The Life of Reason. Dover Publications Inc. 1980. Vol 1. Chapter 12, page 291.
2. A.N. Whitehead. Science and the Modern World. Free Press Paperpack. 1925. p.48.

Monday, February 17, 2014

Seven at a single blow

I still have lunches with my old chums from my days at JPL. On several occasions I've heard, "OK, we know there's problems. But what would you do to make things better?"

The question has been nagging at me. It's one thing to sit on the sidelines an take pot shots, it's quite another to take specific actions. I've been thinking...

How would you even approach the task of changing a very large, inflexible government bureaucracy to refocus its energies from a stovepiped, hardware-centric mindset to a unified software/systems-centric approach? New Leadership? A massive reorganization? A huge infusion of cash? Nah. It's all been tried with no more impact than a sand flea's assault on an elephant. Why?

Large bureaucracies are a civilization's great stabilizing force. The longest-lived civilizations, Egyptian, Chinese, Roman1, all developed powerful bureaucratic institutions. They arise naturally around the seat of power and grow into competitive dominions that will perish if not defended. Since change threatens the division of authority, the power of its management, the livelihood of its minions, and even the bureaucracy's very existence, protective barriers emerge like great jetties against the tide of change. So while bureaucracies stabilize, they also paralyze. It is their nature. Any attempt to change the bureaucracy is a non starter.

What's needed is a strategy that would produce a change by using a bureaucracy's existing machinery.

Grimm O pequeno alfaiatezinho02There's no explaining why, but the story of the Valiant Little Tailor sprang to mind. If you don't remember, the little tailor is an insignificant person who was about to enjoy a lunch of bread and jam when, to his dismay, a swarm of flies lands on his jam. In a moment of pique, he grabs a piece of cloth and swats the flies. The result: 7 flies meet their maker. The little tailor takes great pride in this feat. He wants the world to know, so, he embroiders a sash with the words "seven at one blow" and sets off to make his fortune and eventually rule a kingdom. 3

The story got me thinking. What small feat might trigger the dominos of change, get the Agency out of the current rut and put the space program on a path to develop a new space system that would excite public interest and reestablish the nation's leadership in space technologies? What would disrupt the division of bureaucratic authority and trigger the survival reflex. What if there was a challenge comparable to that faced by the engineers in the Apollo program. What if the Agency decided it was going to build a new kind of system for a flagship mission with requirements that simply could not be produced by current practice? What if there were 7 requirements that would set in motion a fundamentally constructive change?!

"...almost all really new ideas have a certain aspect of
foolishness when they are first produced."4 A.N. Whitehead
Inspired by this challenge, I put in a good bit of careful procrastination, day dreaming, weed pulling and cloud watching. Failing any real insight or inspiration I finally sat down and dashed off 7 requirements that, if mandated, might force the development of a new type of space system--one that would set a new standard for how NASA conceived and developed its spacecraft.
Note: These are non-functional requirements; i.e. they describe how the system would be built or how it would perform. (By contrast, functional requirements describe what functions the system must perform.)

The 7 Requirements:

  1. The proposed mission budget shall account for the mission's complete cost-of-ownership from formulation through operations where operations includes an extended mission phase that is three times the duration of the primary mission.
  2. Rationale: Mission proposals typically reflect a minimum operations budget as an artificial means of reducing the cost to fit under a cap. The practice is always successful because, once the systems is flying, it becomes a valuable asset with vested interests that no seasoned bureaucratic would abandon. Consequently, extended missions are almost always funded. However, since the actual cost of lifetime operations is never reckoned in the original proposal, there is no justification for innovations that would make a spacecraft more reliable or cheaper to operate. If the full cost of ownership was taken into account, that balance would shift and it would be cost-effective to build a better spacecraft.
  3. The flight and ground systems shall be unified into a architecture with a common design philosophy and shared information architecture
  4. Rationale: A common flight/ground architecture would undo the existing stovepiped architecture that's been frozen in place by the vested institutional factions that comprise the bureaucracy. The shared information architecture would narrow the gap between systems and software engineering and eliminate the class of errors that lead to the loss of the Mars Climate Orbiter and Mars Polar Lander.5
  5. The project shall be limited to 36 months.
  6. Rationale: The products created in initial phase of the project are usually cast off when the 'real work' begins. A short schedule will force the team to act decisively from the onset. And, because the schedule is short, institutions will be motivated to invest new approaches before the mission development begins.
  7. (a)The technical lead shall include a 'decide-by' date for every technical decision requiring management approval. (b)Managers shall respond to these decision requests with an unambiguous decision before the 'decide-by date' or be replaced.
  8. Rationale: In a large bureaucracy, decisions are inherently risky. Managers tends to delay decisions until the options are no longer relevant or until a committee consensus has been reached. This requirement would encourage managers to assume individual responsibility. It would also provide the development teams with timely approvals.
  9. Each functional requirement shall include at least one test scenario with success criteria.
  10. Rationale: Requirements expressed as natural language shall statements are typically vague. Even modeled requirements may leave plenty of room for interpretation. By comparison, a test scenario clarifies the intent of a requirement by providing a concrete and specific instance of what the function should do. Clearer requirements help programmers build the right product and avoid costly defects from miscommunication.
  11. The staff needed to complete all mission operations activities shall be limited to 5 engineers.
  12. Rationale: Operations for a flagship mission like MSL typically require dozens of engineers and an annual budget of double-digit or triple-digit megabucks. A team this size is needed to manage trajectory change maneuvers, in situ maneuvers, science planning and recovery from system faults. However, these costs could be dramatically reduced if the system was conceived and built as a unified flight-ground system designed for autonomous flight. Without a requirement to fly with a very small team, the Agency business model coupled with the need to feed the current engineering cadre will prevent the development of an autonomous system.
  13. Board members for software reviews shall be prepared to provide the development team with a detailed technical accounts of any reported findings.
  14. Rationale: In the current process-centric, Agency culture, review boards are inclined to focus on checklists of documents without close examination of the engineering described in those documents. Board members seldom have the time or skills necessary to examine key technical design choices. If board members were knew they might have to discuss technical details, the board would prefer in-depth, engineering-focused presentations to a rehearsal of process compliance.
No doubt someone with superior engineering judgment would produce a better list. However, even this paltry list could serve to shake things up. But only if someone with authority was foolish enough to confront the giants and unicorns and insist on building a system with 7 disruptive requirements. It's an unlikely prospect— bureaucratic managers are selected because they lack those very qualities. However, if by some oversight or error, the highest levels of the Agency approved a disruptive project, I'd walk away from the joys of retirement and do my darndest to secure a spot on that team.


1. The Egyptian empire lasted roughly 5,000 years. The Chinese empire roughly 3,500 years. The combined Roman State (Republic + Empire) lasted about 1,100 years.
2. Interestingly both Roman and Chinese bureaucracies were staffed by eunuchs.
3. Fortunate he was! The little tailor went on to outwit 3 giants, a wild boar and a unicorn on the way to win a princess and a kingdom. In this story, hubris pays off.
4. Whitehead, A.N., Science an the Modern World. Free Press Paperback. 1925. p.47.
5. I know of one architectural approach that meets this requirement: The Mission Data System (MDS). MDS was developed in the last decade, but never used. Perhaps there are other approaches.


Friday, January 31, 2014

The Maintenance Mindset

From Ratus rattus: A digression from the previous post

Excerpt

Fiat Lux Canticle map
A map of North America in 3174
from "A Canticle for Leibowitz."
Software development and software maintenance are fundamentally the same activity and can be funded and managed the same."

From A Canticle for Leibowitz1

Excerpt
For Man was a culture-bearer as well as a soul-bearer, but his cultures were not immortal and they could die with a race or an age
It was a busy and satisfying holiday season. Blogging took a back seat to out-of-town guests and holidays events.

Not all holiday surprises are good. The warp and woof of home repair marches on no matter what the season. 2013 departed with a few unanticipated, and costly, home maintenance levies. So maintenance, in particular maintenance of software systems, has been on my mind.

In the Ratus rattus posting, I listed a few assumptions about the "immutable" facts of life that adversely impact the lives of the NASA software development community—assumptions that must be challenged if we are to build the next generation space system. The list included the claim that NASA management approaches development and maintenance activities as if they are fundamentally the same.

To those who develop software for a living, the distinction between development and maintenance activities may seem obvious. Both entail requirement development, design, coding and test. However, in my experience as a development manager, the differences are fuzzy and easily misunderstood—especially by NASA senior managers and executives who have no experience as programmers. The misunderstanding leads to unrealistic schedules and budgets and, ultimately, to the miseries described 40 years ago in the Mythical Man-Month.

The skeptics among you will have doubts about that claim. In this post, I'll try to explain why this is a mountain and not a molehill. First I'll distinguish how software development differs from software maintenance. Then I'll discuss how the 'maintenance mindset' has become woven in to the fabric of the agency. At the risk of being sketchy, I'll try to be brief.

First a quick comparison.
Availability A system under development has never been deployed and is used only by programmers and testers.
A system under maintenance has been deployed and is being used by mission customers.
Code Change During development, a code change can be introduced without affecting a mission customer.
In maintenance, a change may affect a mission customer, and the impact of that change must be studied before the change is implemented and the system is redeployed.
Requirements During development requirements for the product change as the team balances known-customer needs and discovered-costumer needs against the realities of schedule a budget.
During maintenance, the requirements (or more likely a subset of the requirements) have been implemented, and requirement changes are limited to correcting errors and adding new functions that address customer needs.
Design During development the system design quickly evolves as the team discovers how to address changing requirements.
During maintenance, the system design exists and design changes are confined to correcting defects or fitting in new structures that come with new functions.
Architecture During development the system architecture is evolving.
During maintenance, the system architecture is fixed.
Interfaces During development interfaces morph as the team works through the intricacies of getting the new code to work with existing code.
During maintenance, interfaces become brittle with time and changes may have catastrophic consequences.2
Test During development the test regimen, like the code it tests, is in flux as new tests are created and older tests break.
During maintenance, the test regimen is established and serves as a standard of system readiness.
Programming to Test Ratio During development the programming budget is typically 100-400 percent of the integration-and-test budget.3
During maintenance the programming budget is typically less than 10% of the integration-and-test budget. After all, the product is 'done.'

To sum up, changing a system under maintenance is expensive; changing a system under development is cheap—at least relatively cheap. A poorly planned change to a system under maintenance will likely produce unintended and expensive consequences. So, it is quite sensible that, despite the considerable added expense, changes to a system under maintenance should be subject to elevated standards of governance. It is common practice in NASA that a change to a deployed (i.e. maintained) system requires mountains of paper work and the blessing of one or more change boards.

Conversely, a relaxed standard of governance is appropriate for a system under development—otherwise development costs significantly inflate. Nothing is inherently wrong with the added cost, so long as budgets and schedules match. When they don't, the stage is set for failures that even the best software engineers can not overcome. However, despite the shrinking budgets, it has become common practice in NASA to apply the same governance approach to both maintenance and development tasks. The result: it has become nearly impossible to make fundamental improvements to our systems.

Why has this happened? It's the natural outcome of a historic trend. Think back.

The great innovations in NASA took place when Americans and Soviets were locked in a competition for technical ideological ascendency. It was the era of Apollo and Voyager. The Nation's prestige was on the line. The Agency was well-funded. Engineers solved problems for the first time. Real success was required. Mistakes were expected. Reasonable risks were embraced. 'New' was necessary.

A decade rolls by. Astronauts walk on the moon. Jupiter and Saturn have their close ups. But, with the national prestige assured, the public interest in wanes. Space is no longer prime time—coverage of the Moon landings disappear below the fold. The budgets shrink. And, there are technical problems in paradise. NASA is no longer perfect. The Shuttle Program is delayed two years while engineers burn the night oil trying figure out how to safely attach the heat shield tiles.4 Congress nearly cancels the shuttle program. Meanwhile, flagship science missions like Galileo are dogged by technical and budget problems. Things were not going smoothly. 'New' was becoming risky.

Another decade passes. There are more problems. The big science missions are overrunning schedule and budget. The space station project is stuck in a mire of international politics and technical indecision. Worst of all, there's is a major accident. Meanwhile, the Agency's budget is shrinking. The priorities shift to saving cost and risk reduction. NASA's engineering culture morphs. In the Manned programs, the engineering talent is outsourced. NASA engineers no longer build new systems, rather, they oversee the work of the contractors who build systems that are optimized for cost at the expense of innovation. At the same time, project managers shun new design concepts for science spacecraft because cloning of the previous mission is perceived as being less risky and more cost effective.

It's now 2014 and it's been decades since NASA developed a novel system.5 In particular, the engineering techniques, the avionics, the operational concepts and, above all, the design of the integrated software are fundamentally unchanged. In essence NASA's contemporary approach to space system development amounts to building a minor variant of the last mission.6 That would seem to be the cheapest least risky thing. But consider this...

Each new variant conveys cruft from the previous missions. As the cruft accumulates the software grows brittle and breaks in unexpected ways. The more brittle the software, the more it costs to maintain.

How exactly does software grow brittle? Consider the software needed for a robotic planetary explorer.

N2 Chart Key Features
As the number of inter-dependencies grow
the greater the chance that a change will break the system
An entire system includes a LOT of software. For example: There is software needed for mission design. There is software for operations. There is software for capturing, displaying and storing telemetry. There is software for commanding. There 's software for communications. There is software for turning the data into pictures. That's not to mention the onboard flight code, dozens of other domain specific tools, thousands of tests and thousands of user scripts. All told, 20-30 million lines of code is needed to build and fly a robotic planetary explorer. That's a lot. And, it all has to work together.

Here's the rub: no one really knows exactly how all those pieces fit together. No one knows the dependencies. No one knows what software will break when a change is made. The longer a system has been around, the less is known, the more brittle it becomes. For example, a recent operations system upgrade to the Cassini ground system took over a year and multi-millions of dollars to complete. That was just the OS! As a practical matter, managers responsible for older systems will minimize code changes and invest in as much testing as schedule and budget permit.

With each iteration of the old systems, the problem is multiplied. There's s a dilemma. Do you fund maintenance for each version of each application for each mission or do you fund the maintenance of a single version that's used on several missions. The latter is the rule because superficially it is cheaper to maintain a single version. However, that means that every change now impacts all users and maintenance cost sky rocket. That's why 'New' is perceived forbiddingly expensive and risky.

After 3 decades of rebuilding the same systems, the Agency has lost the instincts for new development. There is a single, standard for the governing software tasks. Developers of new products are expected to produce project plans, requirements and design documents long before the system is understood well enough to produce those artifacts. And, they must do that on unrealistically low budgets that were formulated to compete with specious cost and risk savings assumed for reuse from last mission. In other words, the obligations are going up while the budgets are going down. The cards are stacked against the developer before the project begins.

Perhaps the most worrisome consequence of this maintenance mindset is the lost ability to distinguish the essential from the nonessential. The original reasons for many interfaces, applications, and requirements have been forgotten. Perhaps a file indexing tool or an identifier in an data structure or a data conversion format with its suite of conversion tools were once needed. In time these artifacts become 'necessary' and are sustained with dedicated budgets, managers and staff. Reduced cost and risk aside, any new approach that dispenses with one of these 'necessities' will confront a fierce political and technical struggle from competent people whose jobs are on the line. The fight typically consumes enough of the development resources to renew the skeptics belief that improvements are bound to fail.

This is admittedly a grim picture, but a natural one for a mature bureaucracy like NASA. Bureaucracies will develop organization structures, institutional policies and a management cadre that resists change. It is a natural as the Fall following Summer and Spring. There is vitality in the private sector; corporations who do not evolve, perish. Not so with government bureaucracies. They are especially resistant to change and tend to persist so long as there's an influential constituency on the receiving end of some benefit. Unless there's a crisis of dynamic proportions, no one can expect a government bureaucracy to change.

One the bright side, there are examples of organizations who have found ways around the stagnation. The Lockheed Skunk Works is the premier example. The secret: provide a talented group with a high degree of autonomy. The talent exists in NASA. Now if only there was an imaginative and gutsy executive who could live with the autonomy.

Does anyone see an asteroid headed in our direction?



1. Miller, W. M., "Canticle for Leibowitz." J.B. Lippincott & Co. 1959.
2. Mission users create hundreds, if not thousands, of special purpose applications using both official and unofficial interfaces. Few of these are visible to the development team. Any interface change to a deployed system may break significant parts of a working system in unpredictable ways. This 'brittleness' gets worse with time.
3. Depending on how many times the development team has build similar systems.
4. Williamson, R.A. "Developing the Space Shuttle." from Exploring the Unknown, Chapter 2. P.176. http://history.nasa.gov/SP-4407/vol4/cover.pdf
5. The Constellation Program was no exception. During my stint on the Program, I saw heard upper management sincerely proclaim the intention to build something new. Initially, everyone I knew working on the project was full of hope that we could at last build a system that embodied what we knew needed to be done. However, in the end, the constraints of budget and schedule, the necessity of an affordable bid from a contractor, and the bureaucratic drive for consensus lead to a software system that a mere rehash of all that was built before.
6. Any improvement is cause for great concern. During my last assignment, the decision sue a commercial database for processing and storing telemetry data caused great handwringing. After a year of discussion, the technical leadership remained undecided. Meanwhile, a new mission was seriously considering using a telemetry processor that 15 years out of date.



Saturday, December 21, 2013

Turning the system inside out

From Ratus rattus: A digression from the previous post

Excerpt
"Sponsors and upper management should not be exposed to development details even when those details drive cost and schedule." 
Santi di Tito - Niccolo Machiavelli's portrait
Niccolo Machiavelli (1469-1527)

From The Prince1

Excerpt
There are thee different kinds of brains, the one understands things unassisted, the other understands things when shown by others, the third understands neither alone nor with the explanations of others. The first kind is most excellent, the second also excellent, but the third useless. (Chapter 22, page 104)


In the Rattus rattus posting, I listed a few "frustrated utterances of immutable facts" that adversely impact the lives of the NASA software development community.

The single greatest challenge I faced as a NASA software development manager was finding ways to communicate key decisions to management without delving into technical details. That's right. It was inevitably a mistake to get technical with our management. If we did, we would likely be derailed by tangential questions or hostile interlocutors.

This was a fact of life I never fully accepted. After all, I was working for NASA, the home of advanced technologies, best and the brightest, the nation's stake in the future.

It was 1998. I was fresh blood, full of ideas; the "ancestral pieties"2 didn't apply. What I saw was very smart people working with decade-old technologies. Time to put the past behind. I'd been hired into a group with forward thinkers in leadership roles. We wanted to use the new tools: object-oriented languages, real-time operating systems with protected memory and compilers that supported generic programming. I was about take part in building the next generation space system using state of the art software.

Our goals were at odds with a lunch-time scuttlebutt that was punctuated with aphorisms like "software is an evil necessity," or "there's never time to do it right; there's always time to do it over." As yet, I had no appreciation of the machinery that preserved the established order. I would get my first exposure soon enough.

The implementation phase of our project was about to start. The first major review was around the corner.3 The team had decided to adopt C++ as our programing language. We needed funds for new tools, infrastructure and training. We needed management buy-in. Our manager asked me to present a rationale for using a new language and not selecting Ada or C, the programming languages used on the last missions.4 I prepared a balanced, 20-slide deck with code examples that illustrating the benefits and the pitfalls of C++.5

Five slides into the pre-review walk-through, my manager sent me to the showers. I had too much detail. I had highlighted potential difficulties. By discussing issues that would interest a responsible software developer, I had unwittingly painted a picture of a disaster in the making. My boss warned me that my material would freak the project management and we would surely get 'help' we did not want. If you happen to be a software engineer working in a hardware-centric, government-sponsored bureaucracy, the last thing you need is 'help.'

I was learning. Slowly. There were bigger surprises ahead.

Since I was the new guy, I often reached out to the team's most-experienced programmers for advice. We were about to start programming, but I did not yet have adequate requirements. One day, over coffee, I started grousing to one of the senior guys about the lack of requirements. He smiled knowingly. "Our requirements were useless," he said. By requirements he meant "shall" statements. Given what I'd seen, I had to agree. My personal favorite awful requirement was, "The software shall not harm the hardware." "The Systems Engineers haven't a clue," he said, "we just have to figure it out ourselves."

"What about testing?" I asked. You need requirements to know what to test. "We test it," he said. He meant the programming team. "We show the testers what we did and they just repeat it. No value added." Suffice to say it is NOT considered a best practice for programmers to test their own code. Then there was this clincher. "All that counts is the code. If it's in the code, it's on the mission."

So happened that this particular programmer was an expert user of the system he was building. He knew what the system should do and could build a usable system without requirements. Still, I was skeptical that his code pass the delivery review. Surely there would be a reckoning. There wasn't. That review went something like this: "Code delivered on time and tested." His delivery was a rip-roaring success. Management was sastified--there was no apparent cause for worry.

During the past decade, our development practices became more rigorous. The agency adopted a set of required processes for software.6

In spirit, these process mandates are reasonable. In practice, they levy a significant, typically unfunded, burden that produces a mountain of paper that describes the code and how it was built. So much paper that only a small portion of the documents are carefully read. An even smaller portion are treated to thoughtful analysis by an engineer with sufficient expertise to render a useful opinion. Nevertheless, these documents become an official record of engineering thoroughness--a certification of the quality of the code. When reviews roll around, management can conveniently meet their obligations by ticking down a list of required documents to see if any of their number is missing.

It's a very practical arrangement that has become settled convention. Developers are free to do their work without exposing coding details or any risks that might be associated with design choices. Managers are assured by the process machinery that all is in order without taking the trouble to understanding the software. For much of the schedule, the project purrs along with happy sounding Earned Value metrics until the predicable budget overruns and schedule slips (which the required process did nothing to alleviate) light up the FEVER charts with red and yellow like a Christmas tree .

For years I failed to abide by this unwillingness to understand the details of the software. However, as I assumed greater management responsibility, I came to appreciate, even accept, why software engineering was the red-headed step child on NASA projects. In the conventional view, a spacecraft is fundamentally a very complicated piece of hardware; it just happens to have some software inside. Managing a $400-500M enterprise is hard enough without getting bogged down in the minutia of software piece parts. The spacecraft-as-hardware is a cultural mindset with roots that reach back to the Apollo era. The vestiges from that time live on in the project WBS and major milestone reviews. For example: a typical WBS places the flight software under the avionics subsystem a bureaucracy away from the ground system software. Similarly, a typical three-day, 36-hour gateway review, allows but a couple of hours for the discussion of the flight and ground software efforts.

And yet, in project after project, Brooks' 40-year-old admonitions reign supreme--software remains a persistently vexing management problem. The code is late and over budget. It doesn't do what it is supposed to do. There are bugs. The maintenance costs a fortune. There always a plague of technical gotyas that resist a simple fix. The tools that worked for the last system no longer work. The explanations from the software people are arcane and bewildering. Is it any wonder why project management is distrustful of software when such a small portion of the budget repeatedly causes so much trouble?

I have been on the receiving end of this management skepticism. There is no good answer to questions like: "Why are you reinventing the wheel?" Or, "Are those changes really needed?" Technical reasons, no matter how good, sound defensive or seem to obfuscate. I've heard it on reliable authority that in their corner offices senior managers confide to each other that the software problem stems from a lack of discipline and a lackadaisical attitude about commitment. So when a crisis of budget or schedule beckons and a management decision is required, it's usually rendered with the preamble, "I don't know anything about software, but..." True enough. It's a decision made on the basis of mistrust without knowledge of the details.

If you've read elsewhere in this blog, you'll know I believe that the next-generation space system must be a software-intensive system and not a spacecraft-as-hardware system. In other words, the design and implementation of the software would become an overarching project concern that links power, propulsion, mass, attitude control, navigation, operational concepts and fault protection. This means turning the system concept inside out so that project leadership makes a priority of understanding the software and how it connects across the system. Until that happens, the development of a smart, reliable, affordable system capable of complex operations, human or robotic, will remain beyond the reach of space system engineering.

Still, the underlying problem of managing a large, complicated development effort remains. The leadership must be able to understand the software and still orchestrate the work of the many engineering efforts by the collaborating disciplines. No one can master all.

Management will need to have an intuition about the software to grasp which details matter. Intuition that only comes from the experience of writing code under deadline pressure for an unknown user. Code that is designed for change, reuse and longevity. The kind of code Brooks called a "programming systems product." Only then will a manager have the gut-wrenching experiences needed to understand why software development is not like a music box. It is a discovery process that varies with the maturity of the team, the tools and the product. Failing that experience, it's very unlikely a manager will have a reliable intuition.

To the best of my knowledge, there is not, nor has there been, a single senior manager in NASA who has worked as a professional programmer. After all, NASA is an mature, hardware-centric, government bureaucracy with an entrenched culture. Cultural adjustments are disruptive. Of all the challenges that face the Agency, introducing a software-centric focus may be the most difficult.

Whitehead famously writes about the advance of civilization through the effect of certain ideas. "An idea is a prophecy which procures its own fulfillment."7 A reworking of NASA to prepare for the development of a next generation's space system is not beyond the realm of possibility. The transition could be realized in a single administration by an enlightened, determined leadership. It has happened in the past when the agency was born. It could happen again.


1. Michiavelli, N., The Prince. Translated by Ricci, L. Revised by Vincent, E.R.P., Oxford University Press, World Classics. Reprinted 1968.
2. Nifty phrase lifted from A.N. Whitehead. "Adventures in Ideas". Free Press Paperback. 1933.
3. A Preliminary Design Review (PDR)that occurred in the spring of 1998.
4. The Cassini flight software was written in Ada. The flight software for Pathfinder was written in C.
5. C++ is a very powerful but difficult programming language because it is easy to make subtle errors that lead to bugs and performance issues. I personally had a dozen books well-read books that provided programming guidance.
6. For a sample of the required NASA processes see the NASA Process web site.
7. A.N. Whitehead. "Adventures in Ideas". Free Press Paperback. 1933. Page 42.