Showing posts with label data analytics. Show all posts
Showing posts with label data analytics. Show all posts

Tuesday, September 27, 2016

New Book: Business Intelligence with R by Dwight Berry

If you are new to data science and learning the R language, let me recommend this new gem of a book, Business Intelligence with R, by my friendr (the term I just coined to describe R users who help each other), Dwight Barry: https://leanpub.com/businessintelligencewithr

Business Intelligence with R serves as a great cookbook that can save you hours of frustration learning how to get the basics going. Even if you're an old pro, the book serves as a handy desk reference.

Also, please consider the personal note that Dwight sent to all of his beta readers:
Perhaps most importantly, I've also decided to give all proceeds to the Agape Girls Junior Guild, which is a group of middle-school girls who do fundraising for mitochondrial disorder research at Seattle Children's Research Institute and Seattle Children's Hospital. While the minimum price for this book will always be free, if you're the type who likes to "buy the author a coffee," know that your donation is supporting a better cause than my already out-of-control coffee habit. :-)
Business Intelligence with R serves a greater cause.

Wednesday, January 07, 2015

An Interesting Christmas Gift

Over the holidays, the New York Times delivered an unusual juxtaposition of headlines and content, and apparent lack of self-awareness, to illicit such a hearty chuckle from its readers as to make the cheerful Old Saint jealous.


[image originally provided by @ddmeyer on Twitter]

To those imbued with the skill of basic high school Algebra 1, the information in the article about Sony’s revenues for the first four days of release of “The Interview” were enough to solve a unit value problem. If we let R = the number of rentals, and S = the number of sales; then,
  • R + S = 2 million 
  • $6*R + $15*S = $15 million 
With a little quick symbolic manipulation, we see that S = 1/3 million in sales and R = 5/3 million in rentals. That exercise provided just enough mental stimulation and smug self-righteousness to prepare for the day’s sudoku and crossword puzzles. #smug #math

However, not too far into the sudoku puzzle we might realize that a deeper, more instructive problem exists here, a problem that actually permeates all of our daily lives. That problem is related to the precision of the information we have to deal with in planning exercises or, say, garnering market intelligence, etc. A second reading of the article reveals that the sales values, both the total transactions and the total value of them, were reported as approximations. In other words, if the sources at Sony followed some basic rules of rounding, the total number of transactions could range from 1.5 million to 2.4 million, and the total value might range from $14.5 million to $15.4 million. This might not seem like a problem at first consideration. After all, 2 million is in the middleish of its rounding range as is $15 million. Certainly the actual values determined by the simple algebra above point to a good enough approximate answer. Right? Right?

To see if this true, let’s reassign the formulas above in the following way.
  • R + S = T 
  • $6*R + $15*S = V 
where T = total transactions, and V = total value. Again, with some quick symbolic manipulation, we can get the exactly precise answers for R and T across a range of values for T and V.
  • S = 1/9 * V - 2/3 * T 
  • R = T - S 
Doing this we now notice something quite at odds with our intuition - the range of variation between the sales and rentals can be quite large as we see in this scatter plot:



[Fig. 1: The distribution of total transaction values for various combinations of rental and direct sales numbers.]

Here we see that the rental numbers could range from about 800 thousand to 2.4 million, while the direct sales could range from nearly 0 to 700 thousand! Maybe more instructive is to consider the range of the ratio of the rentals to direct sales:


[Fig. 2: The distribution of the ratio of rentals to direct sales for various combinations of rental and direct sales numbers.]

If we blithely assume that the reported values of sales were precise enough to support believing that the actual value of rentals and unit sales were close to our initial result, we could be astoundingly wrong. The range of this ratio could run from about 1.11 (for 1.5 million in total transactions; 15.4 million in sales) to 215 (for 2.4 million in total transactions; 14.5 million in sales). If we were trying to glean market intelligence from these numbers on which to base our own operational or marketing activities, we would face quite a conundrum. What’s the best estimate to use?
Fortunately, we can turn to probabilisitic reasoning to help us out. Let’s say we consult a subject matter expert (SME) who gives us a calibrated range and distribution for the sales assumptions such that the range of each distribution stays mostly within the rounding range we specify.

[Fig. 2a, b: The hypothetical distribution of the (a) total sales transactions and (b) total value assessed by our SME.]

Using the sample values underlying these distributions in our last set of formulas, we observe that in all likelihood - an 80th percentile likelihood – the actual ratio of the rentals to sales falls in a much narrower range – the range of 3 to 9, not 1.11 to 215.

[Fig. 3: The 80th percentile prediction interval for the ratio of the rentals to sales falls in the range of 3 to 9.]

Our manager may push back on this by saying that our SME doesn’t really have the credibility to use the distributions assessed above. She asks, "What if we stick with maximal uncertainty within the range?” In other words, what if, instead of assessing a central tendency around the reported values with declining tails on each side, we assume there is a uniform distribution along the range of sales values (i.e., each value is equally probable to all values in the range)?



[Fig. 4a, b: We replace our SME supplied distribution for (a) total sales transactions and (b) total value with one that admits an insufficient reason to suspect that any value in our range is more likely than any other.]

What is the result? Well, we see that even with the assumption of maximal uncertainty, while the most likely range expands by a factor of 2.7 (i.e., the range expanded from 3-9 to 1.7-18), it still remains within a manageable range as the extreme edge cases are ruled out, not as impossible but as fairly unlikely.

[Fig. 5: Replacing our original SME distributions that had peaks with uniform distributions flattens out the distribution of our ratio of rentals to sales, causing the 80th percentile prediction interval to widen. The new range runs from about 1.7 to 18.]

The following graph displays the full range of sales and rental variation that is possible depending on our degrees of belief (as represented by our choice of distribution) about the range of total transactions and total value.

[Fig. 6: A scatter plot that demonstrates the distribution of direct sales and rental combinations as conditioned by our choice of distribution type.]

By focusing on the 80th percentile range of outcomes in the ratio of rentals to sales, we can significantly improve the credible range to estimate the rentals and direct sales from the approximate information we were given.

[Fig. 7: A scatter plot that demonstrates the distribution of direct sales and rental combinations as conditioned by our choice of distribution type, constrained only to those values in the 80th percentile prediction interval.]

Precise? Not within a hair’s breadth, no, but the degree of precision we obtain by employing probabilities (as opposed to relying on just a best guess with no understanding of the implications of the range of the assumptions) into our analysis improves by a factor of 13.1 (assuming maximum uncertainty) to 35.2 (trusting our SME). If our own planning depends on an understanding of this sales ratio, we can exercise more prudence in the effective allocation of the resources required to address it. Now, when our manager asks, “How do you know the actual values aren’t near the edge cases?”, we can respond by saying that we don’t know precisely, but using simple algebra combined with probabilities dictates that the actual values most likely are not.

Thursday, July 17, 2014

When A Picture is Worth √1000 Words

This morning @WSJ posted a link to the story about Microsoft’s announcement of its plans to lay off 18,000 employees. This picture (as captured on my iPhone)...

[click image to enlarge]

...accompanied the tweet, which is presumably available through their paywall link.

While I’m really sorry to hear about the Microsoft employees who will be losing their jobs, I am simply outraged at the miscommunication in the pictured graph. (This news appeared to me first on Twitter, and the seemingly typical response on Twitter is hyperbolic outrage.)

Here’s the problem as I see it: the graph communicates one-dimensional information with two-dimensional images. By doing so, it distorts the actual intensity of the information the reporters are supposed to be conveying in an unbiased manner. In fact, it makes the relationships discussed appear much less dramatic than it actually is.

For example, look at Microsoft’s (MSFT) revenue per employee compared to Apple’s (AAPL). WSJ reports MSFT is $786,400/person; APPL, $2,128,400. The former is 37% of the latter. But for some reason, WSJ communicates the intensity with an area, a two-dimensional measure, whereas intensity is one-dimensional. Our eyes are pulled to view the length of the side of the square as a proxy for the measurement being communicated. The sides of the squares are proportionally equal to √(786,400) and √(2,128,400); therefore, the sides of the squares visually communicate the ratio of the productivity of MSFT:AAPL as 61%. In other words, the chart visually overstates the relative productivity of MSFT's employees compared to that of AAPL's by a factor of 1.62.

If the numbers are confusing there, consider this simpler example. The speed of your car as measured by your speedometer is an intensity. It’s one dimensional. It tells you how many miles (or kilometers, if you’re from most anywhere else outside the US) you can cover in one hour if your car maintains a constant speed. Your speedometer aptly uses a needle to point to the current intensity as a single number. It does not use a square area to communicate your speed. If it did, 60 miles per hour would  look 1.41 times faster than 30 miles per hour instead of the actual 2 times faster that it really is. The reason for this is that the the sides of the squares used to display speed would have to be proportional to the square roots of the speed. The square roots of 60 and 30 are 7.75 and 5.48, respectively.

For your own personal edification, I have corrected the WSJ graph here:

[click image to enlarge]

Do you see, now, how much more dramatic the AAPL employees' productivity is over that of MSFT's?

This may not seem like a big deal to you at the moment, but consider how much quantitative information we communicate graphically. The reason is that, as the cliché goes, a picture is figuratively worth a thousand words. I firmly believe graphical displays of information are powerful methods of communication, and a large part of my professional practice revolves around accurately and succinctly communicating complex analysis in a manner that decision makers can easily consume and digest. But I’m also keenly aware of how analyst and reporters often miscommunicate important information via visual displays, either by design, inexperience, or by trying to be too clever. I see these transgressions all the time in the analyses I’m asked to audit.

The way we communicate information is not just a matter of style for business reporters. We often make prodigious decisions based on information. If information is communicated in a way that distorts the underlying relationships involved, we risk making serious misallocations of scarce resources. This affects every aspect of the nature of our wealth - money, time, and quality of life. The way we communicate information bears fiduciary responsibilities.

For discussion sake I ask,

  1. How often have you seen, and maybe even been victimized by, graphical information that miscommunicates important underlying relationships and patterns?
  2. How often have you possibly incorporated ineffective means of graphically communicating important information? (Pie charts, anyone?)

If you want to learn more about the best ways to communicate through the graphical display of quantitative information, I highly recommend these online resources as a starting point: