24 July 2007
Books I really, really, really want to index
Perhaps because most of my indexing work is on books that are rather typical for nonfiction reference books -- technical titles like Measurement, Analysis, and Control Using JMP and resource guides like Dx/Rx: Colorectal Cancer -- I jump for joy when I get something so off the beaten path that I renew my love for this job. For example, I recently completed the index for First Position, a collection of biographies of ballet dancers; more recently, I indexed Sensual Knits and Sensual Crochet, both beautifully photographed books of designs and patterns.
But now, working in the wee hours of the night, I find myself fantasizing about the books that I really, really, really want to index, books that are just asking to be written so that I, Seth Maislin, can be assigned their indexes. So here's my wish list:
Chihuahuas for Dummies
Don't laugh. You probably have no idea just how far-reaching the Dummies series has become since its long-ago inception as a series for computer use. There's Fantasy Football for Dummies, a book about imaginary sports playing; Stretching for Dummies, a book about limbering up, perhaps in advance of reading Sex for Dummies; Guitar for Dummies, Bass Guitar for Dummies, and the upcoming Rock Guitar for Dummies, which I have to believe compete with each other somehow; and Jewish Cooking for Dummies, a book that, dare I say it, would make me feel guilty to own. Nevertheless, let me make myself clear here. Chihuahuas for Dummies is a real book. I want to index the next edition, you see, because I'm dying to see what changes.
10.9 Seconds: The Joey Chestnut Story
(see http://origin.mercurynews.com/valley/ci_6297731 to get the joke) Yes, this book is my own invention, but the fun part about indexing sports books is that they are so completely self-reverential. (Yes, reverential, not referential.) Written by sports geeks for sports geeks, the authors' language captures the awe-hubris-humor combination achieved by fans and record-breakers when it comes to the sport that is most of their life. It doesn't matter what the sport is, either, so I'm all for those esoteric things like Ultimate Frisbee (I was offered such a book once) and so on. I recently indexed the comprehensive Chasing the Hunter's Dream, a directory of hunting opportunities around the world. This book included both descriptions of "dream hunts" -- think lion hunts in Africa -- and an entire section in the back dedicated to recipes, including a few meals for squirrels -- I mean, of squirrels. And I mentioned First Position in my intro, where at times I felt like I was reading an artist's diary.
How to Work My Body: A Manual
There are a number of sex books out there -- including one for Dummies -- and most of them have indexes. I just finished indexing Him and Her, short and photograph-filled manuals of the sexes, along with instructions to make them work. And I do mean "work": the book about men attempts to explain why they tend not to do chores around the house. (Oh come on, you didn't think I'd use an erotic example of "work", did you? :-) These books, produced by the same group of people who made Sensual Crochet, were a joy of sex to index, especially once I realized that most of the anatomy-filled books that I index are about abnormal anatomy: prostate disorders (100 Questions and Answers About Prostate Diseases), gunshot wounds (Criminal Investigation, 2nd edition), and the like. And unlike the traditionally polite sex-instruction book, Him and Her are more about the art than the words -- something that, for eunuchs at least, would make the indexing go much faster.
The user manual to anything only cool people own
I had the honor of indexing the user manual to the Class E series Mercedes-Benz automobile. This full-color production was totally awesome; I spent a lot of time trying to convince myself that reading the manual long before the car's official release was as envy-worthy as owning the car itself. (For many months my friends and family joked that I should paid in cars instead of dollars.) I've indexed the manuals to software applications before, but I have more memories from editing the user guide to a long-since-extinct universal remote control ... and I'm talking back when these things were large control panels. So what other cutting-edge production is taking place? I missed indexing the iPhone manual, but maybe someday I'll get to index the field guide for a nasty-looking military weapon.
Instructions to the 1040 Form
Indexing gets so little press, but that doesn't stop me from wanting to index something that's so popular or high-profile that I can't feel proud. I'll never be a household name, but if I had landed that one magical indexing project with the U.S. Internal Revenue Service, my work might have reached every household. They really were looking for someone, at least for a little while. Even the newest Harry Potter book isn't as popular. Which reminds me: is someone out there indexing Rowland's books? If not, there ought to be. The Unauthorized Index of Harry Potter would be a big seller ... despite use of the word index in the title. Move over, back-of-the-book indexing. We're on the cover now.
Okay, I'm starting to drool.
But now, working in the wee hours of the night, I find myself fantasizing about the books that I really, really, really want to index, books that are just asking to be written so that I, Seth Maislin, can be assigned their indexes. So here's my wish list:
Chihuahuas for Dummies
Don't laugh. You probably have no idea just how far-reaching the Dummies series has become since its long-ago inception as a series for computer use. There's Fantasy Football for Dummies, a book about imaginary sports playing; Stretching for Dummies, a book about limbering up, perhaps in advance of reading Sex for Dummies; Guitar for Dummies, Bass Guitar for Dummies, and the upcoming Rock Guitar for Dummies, which I have to believe compete with each other somehow; and Jewish Cooking for Dummies, a book that, dare I say it, would make me feel guilty to own. Nevertheless, let me make myself clear here. Chihuahuas for Dummies is a real book. I want to index the next edition, you see, because I'm dying to see what changes.
10.9 Seconds: The Joey Chestnut Story
(see http://origin.mercurynews.com/valley/ci_6297731 to get the joke) Yes, this book is my own invention, but the fun part about indexing sports books is that they are so completely self-reverential. (Yes, reverential, not referential.) Written by sports geeks for sports geeks, the authors' language captures the awe-hubris-humor combination achieved by fans and record-breakers when it comes to the sport that is most of their life. It doesn't matter what the sport is, either, so I'm all for those esoteric things like Ultimate Frisbee (I was offered such a book once) and so on. I recently indexed the comprehensive Chasing the Hunter's Dream, a directory of hunting opportunities around the world. This book included both descriptions of "dream hunts" -- think lion hunts in Africa -- and an entire section in the back dedicated to recipes, including a few meals for squirrels -- I mean, of squirrels. And I mentioned First Position in my intro, where at times I felt like I was reading an artist's diary.
How to Work My Body: A Manual
There are a number of sex books out there -- including one for Dummies -- and most of them have indexes. I just finished indexing Him and Her, short and photograph-filled manuals of the sexes, along with instructions to make them work. And I do mean "work": the book about men attempts to explain why they tend not to do chores around the house. (Oh come on, you didn't think I'd use an erotic example of "work", did you? :-) These books, produced by the same group of people who made Sensual Crochet, were a joy of sex to index, especially once I realized that most of the anatomy-filled books that I index are about abnormal anatomy: prostate disorders (100 Questions and Answers About Prostate Diseases), gunshot wounds (Criminal Investigation, 2nd edition), and the like. And unlike the traditionally polite sex-instruction book, Him and Her are more about the art than the words -- something that, for eunuchs at least, would make the indexing go much faster.
The user manual to anything only cool people own
I had the honor of indexing the user manual to the Class E series Mercedes-Benz automobile. This full-color production was totally awesome; I spent a lot of time trying to convince myself that reading the manual long before the car's official release was as envy-worthy as owning the car itself. (For many months my friends and family joked that I should paid in cars instead of dollars.) I've indexed the manuals to software applications before, but I have more memories from editing the user guide to a long-since-extinct universal remote control ... and I'm talking back when these things were large control panels. So what other cutting-edge production is taking place? I missed indexing the iPhone manual, but maybe someday I'll get to index the field guide for a nasty-looking military weapon.
Instructions to the 1040 Form
Indexing gets so little press, but that doesn't stop me from wanting to index something that's so popular or high-profile that I can't feel proud. I'll never be a household name, but if I had landed that one magical indexing project with the U.S. Internal Revenue Service, my work might have reached every household. They really were looking for someone, at least for a little while. Even the newest Harry Potter book isn't as popular. Which reminds me: is someone out there indexing Rowland's books? If not, there ought to be. The Unauthorized Index of Harry Potter would be a big seller ... despite use of the word index in the title. Move over, back-of-the-book indexing. We're on the cover now.
Okay, I'm starting to drool.
Labels: books, fun with indexing
20 January 2007
Foreword to Heather Hedden's upcoming book
I was asked to write the foreword to Heather Hedden's upcoming Indexing Specialties: Web Sites, to be published in 2007 by ITI. Given the importance of this book in the indexing industry, I am reprinting that foreword here. For more information on the book itself (not yet available), visit either ASI's publications page or a list of ITI's indexing publications.
- - - - - -
Foreword
Indexing is not a popular profession by any stretch of the imagination. Not only is it almost completely unknown in lay circles, but let's be honest: writing indexes sounds about as exciting as cleaning the house, but a hundred times harder. Also, if you were born in any year before 1990, the idea of Web indexing sounds like cleaning a house in outer space. I mean, there are no houses in outer space.
The Internet and the Web -- this monstrously huge and growing system of sharing data -- desperately need more information sorcerers like Heather Hedden. Not only does Heather have the talent to recognize when knowledge is missing, but she also has the ability to make that knowledge visible. She starts by learning for herself, and then she loves to share.
Heather and I first crossed paths in my classroom, where I taught a course called "Writing Indexes for Books and Websites." My course was written to explore the questions and theories of indexing, and so couldn't be limited to just books. Heather’s interest went much further, and since then she has explored writing web indexes as a singular discipline. For me, Heather has been a student, an apprentice, and a role model. She's someone I count on to get things done. She has vaulted across the lines from library science to book indexing to web indexing, each time with surprising success, and has since become a renowned and respected expert in the web indexing community.
Indexing Specialties: Web Sites is a book filled with honest, get-it-done advice. Heather is not afraid to talk about the code and the tools, because she has faith in her readers. In her hands, the complicated stuff looks straightforward. Besides, when the technical lessons are over, Heather shows readers how to think about web indexing as well: as a process and as a business. Until now, if book indexers wanted to graduate to the Internet frontier, they had no unified place of reference, no single source of everything they'd want to know. In fact, some of the tools Heather includes in this book were almost completely unknown to indexers until now.
I am excited and pleased to see Heather compiling this knowledge in a book. She has put into print an indexer's Rosetta Stone, which will lead book indexers toward other information management topics like taxonomies, information architecture, and search tools. It's not about complicated coding practices and computer programs, but about the guidelines to getting that A-to-Z index published on the Internet, and doing it right.
She begins by exploring the boundaries of web site indexing, clarifying what kinds of sites need indexing, how they should look, and how they should work. Then she immediately provides the HTML building blocks to making your indexes appear on the Web, the surprisingly simple code you'd need to create index pages, index entries, indentations, hyperlinks, and cross-reference links. If you've never programmed on the Web before and are afraid it's over your head, you’ll be kicking yourself once you see how easy Heather makes it.
Once you're armed with the grammar, you next need the tools to actually write. Heather gives you the detail about the tools (CINDEX, HTML Indexer, HTML/Prep, Macrex, SKY Index Professional, and XRefHT) to create or generate indexes that are ready for web publication. She takes more time exploring the specialized tools of XRefHT and HTML Indexer, two stand-alone web indexing applications, and shows how you can use their features with agility.
The last third of the book is dedicated to the "mindspace" of web indexing. There's more to indexing than just the tools, and so Heather writes carefully about how indexers should approach the job. She addresses the challenges of working out of order, adding anchors, indexing periodicals, and knowing which pages and at what level of detail you should index. She deals in detail with cross-references, language, subentry structure, and format. Finally, Heather dives into the nitty-gritty of the web indexing marketplace, including how to market yourself as a web site indexer.
Web Sites is going to satisfy you immediately and in the long term. On behalf of the American Society of Indexers -- and myself, personally -- I am honored to welcome Heather as an esteemed author in our community.
Seth Maislin
President of the American Society of Indexers (2006-2007)
- - - - - -
Foreword
Indexing is not a popular profession by any stretch of the imagination. Not only is it almost completely unknown in lay circles, but let's be honest: writing indexes sounds about as exciting as cleaning the house, but a hundred times harder. Also, if you were born in any year before 1990, the idea of Web indexing sounds like cleaning a house in outer space. I mean, there are no houses in outer space.
The Internet and the Web -- this monstrously huge and growing system of sharing data -- desperately need more information sorcerers like Heather Hedden. Not only does Heather have the talent to recognize when knowledge is missing, but she also has the ability to make that knowledge visible. She starts by learning for herself, and then she loves to share.
Heather and I first crossed paths in my classroom, where I taught a course called "Writing Indexes for Books and Websites." My course was written to explore the questions and theories of indexing, and so couldn't be limited to just books. Heather’s interest went much further, and since then she has explored writing web indexes as a singular discipline. For me, Heather has been a student, an apprentice, and a role model. She's someone I count on to get things done. She has vaulted across the lines from library science to book indexing to web indexing, each time with surprising success, and has since become a renowned and respected expert in the web indexing community.
Indexing Specialties: Web Sites is a book filled with honest, get-it-done advice. Heather is not afraid to talk about the code and the tools, because she has faith in her readers. In her hands, the complicated stuff looks straightforward. Besides, when the technical lessons are over, Heather shows readers how to think about web indexing as well: as a process and as a business. Until now, if book indexers wanted to graduate to the Internet frontier, they had no unified place of reference, no single source of everything they'd want to know. In fact, some of the tools Heather includes in this book were almost completely unknown to indexers until now.
I am excited and pleased to see Heather compiling this knowledge in a book. She has put into print an indexer's Rosetta Stone, which will lead book indexers toward other information management topics like taxonomies, information architecture, and search tools. It's not about complicated coding practices and computer programs, but about the guidelines to getting that A-to-Z index published on the Internet, and doing it right.
She begins by exploring the boundaries of web site indexing, clarifying what kinds of sites need indexing, how they should look, and how they should work. Then she immediately provides the HTML building blocks to making your indexes appear on the Web, the surprisingly simple code you'd need to create index pages, index entries, indentations, hyperlinks, and cross-reference links. If you've never programmed on the Web before and are afraid it's over your head, you’ll be kicking yourself once you see how easy Heather makes it.
Once you're armed with the grammar, you next need the tools to actually write. Heather gives you the detail about the tools (CINDEX, HTML Indexer, HTML/Prep, Macrex, SKY Index Professional, and XRefHT) to create or generate indexes that are ready for web publication. She takes more time exploring the specialized tools of XRefHT and HTML Indexer, two stand-alone web indexing applications, and shows how you can use their features with agility.
The last third of the book is dedicated to the "mindspace" of web indexing. There's more to indexing than just the tools, and so Heather writes carefully about how indexers should approach the job. She addresses the challenges of working out of order, adding anchors, indexing periodicals, and knowing which pages and at what level of detail you should index. She deals in detail with cross-references, language, subentry structure, and format. Finally, Heather dives into the nitty-gritty of the web indexing marketplace, including how to market yourself as a web site indexer.
Web Sites is going to satisfy you immediately and in the long term. On behalf of the American Society of Indexers -- and myself, personally -- I am honored to welcome Heather as an esteemed author in our community.
Seth Maislin
President of the American Society of Indexers (2006-2007)
Labels: books, web indexing
26 June 2006
With respect to our loved ones
There are many who believe that books -- and by this I refer to the traditional object of bound paper and print -- are sacred objects. Anne Fadiman wrote in Ex Libris that books are often marketed as if they were toasters, and yet remembered as if they were friends. Certainly a block of paper, ink, and glue would not hold such an esteemed place in our hearts if there were nothing transcendent about it. The way in which we archive old books on our shelves because of the memories they inspired in us, whether as a favorite book from early childhood or an intellectual realization from our older years, is not unlike the way a museum places found bones under glass. Unlike the skeletal remains of an ancient animal, however, our old books are neither unique nor unused nor truly old.
The loss of such a book -- from spilled juice, from a disrespectful borrower, from a forgetful moment on a bus -- can be devastating. Almost all titles can be repurchased, in some cases with benefits like a new introduction by the author, an improved detail of scholarly footnotes, or a respectful commentary written with the benefit of time. Rarely, though, is it the corporeal book itself that brings us such pleasure. Had the book been empty, like those ubiquitous writing books and diaries sold at the checkout displays of almost every bookstore, its loss would have gone mostly unnoticed, valued at approximately the retail cost printed on its back cover. No, it is the content that brings us pleasure, with its memories of having been explored.
Content is what makes our books cherished items. Porcelain figures, music boxes, ticket stubs, and toy animals have memories but no inherent content, whereas books can be reopened, reread, and rediscovered. Even when the words don't change -- and they rarely do -- the experience of seeing them with changed eyes and minds is different each time.
Listen, friends, for this is the romantic side to indexing.
As readers -- gentle, voracious, impulsive, or any other adjective that best defines the nature of your reading relationships -- we have an obligation to provide access to these memories, past and future. Even if we do not write, we must endow the writings of others with every tool at our disposal. We cannot guarantee that any single book won't get lost in the attic or destroyed by fire, but we can, as indexers, guarantee that every important sentence within is flagged with accuracy and passion. We can, in the end, turn the writings of others into useful thoughts, moments of learning, and renewable tools for discovery and self-discovery.
The loss of such a book -- from spilled juice, from a disrespectful borrower, from a forgetful moment on a bus -- can be devastating. Almost all titles can be repurchased, in some cases with benefits like a new introduction by the author, an improved detail of scholarly footnotes, or a respectful commentary written with the benefit of time. Rarely, though, is it the corporeal book itself that brings us such pleasure. Had the book been empty, like those ubiquitous writing books and diaries sold at the checkout displays of almost every bookstore, its loss would have gone mostly unnoticed, valued at approximately the retail cost printed on its back cover. No, it is the content that brings us pleasure, with its memories of having been explored.
Content is what makes our books cherished items. Porcelain figures, music boxes, ticket stubs, and toy animals have memories but no inherent content, whereas books can be reopened, reread, and rediscovered. Even when the words don't change -- and they rarely do -- the experience of seeing them with changed eyes and minds is different each time.
Listen, friends, for this is the romantic side to indexing.
As readers -- gentle, voracious, impulsive, or any other adjective that best defines the nature of your reading relationships -- we have an obligation to provide access to these memories, past and future. Even if we do not write, we must endow the writings of others with every tool at our disposal. We cannot guarantee that any single book won't get lost in the attic or destroyed by fire, but we can, as indexers, guarantee that every important sentence within is flagged with accuracy and passion. We can, in the end, turn the writings of others into useful thoughts, moments of learning, and renewable tools for discovery and self-discovery.
Labels: books
14 May 2006
Demands for quantity are misplaced
There is a continuing trend in search engines: more, more, more.
Press releases from Google, like “Google Checks Out Library Books” [December 14, 2004] and “Google Tunes Into TV” [January 25, 2005], hit the media waves in a grand style. For the first time, entire libraries of books, from Harvard and Stanford Universities to the Universities of Michigan and Oxford, and soon the New York City Public Library, will be available from Google’s website. Within limits of copyright, the words of entire books can be searched. Information philosophers are all over this story, essentially declaring that Google will become the public library of the next generation, excited about how the very nature of libraries might change, and scratching their heads over how the book publishing industry is going to survive yet another hit in the market.
In the second, Google (as well as Yahoo!) applauds themselves for once again providing access to a greater diversity of the world’s information, because television’s closed captioning content has been indexed into a Google Video database. Viewers of public broadcasting and basketball are early adopters, and why not? Finally, all those oh-so-deprived sports consumers can satisfy themselves on more than just the videos, statistics databases, press releases, articles, blogs, commentaries, and (don’t forget) live games themselves. Because now they can search among the announcers’ words.
Now when I search for 76ers, instead of getting 1.61 million hits, I’ll get 1.62 million. Phooey.
Our instinctive reaction is to be impressed. I’m thinking of all those Ph.D. theses gathering dust in the Physics-Optics-Astronomy Library at my alma mater, 150-page books without indexes. I’m thinking about Red Sox fans who, for the first time in a very long time, are interested in the World Series.
But I’m also thinking about the catalog search system at my public library, which won’t improve with Google’s additions. For a book already in the catalog, adding its content doesn’t help at all. Instead, we’d be cluttering up the database with a few trillion new words.
So our instincts are wrong.
Why Quantity Hurts
Every few years, search engine companies are finding new ways to promote themselves by bragging about how much they can find. In the early 1990s, Northern Light was independently rated top among competitors because they searched the largest percentage of the World Wide Web: sixteen percent. Time and Newsweek contributors warned, “We’re not finding everything.”
In the late 1990s, articles about the “invisible web” appeared in popular magazines and newspapers, explaining how search engines cataloged only text, image, and sound files, thus skipping over the good content stored as spreadsheets, databases, and fonts. Again came the cry, “We’re not finding everything!”
And now, in just two months, Google and Yahoo have added more to the huge pile of information: library books and closed captioning data. No longer are our searches limited to the billions of files already on the Web. Yay!
It’s all about quantity. Nobody seems to care about quality any more.
Libraries have been struggling to redefine themselves ever since the Internet (and more, the World Wide Web) reached people’s homes. Book publishers also have suffered. The failure isn’t that of these institutions and industries, however, but of the public. The public seems unaware of the natural filtering process inherent in human behavior. Publishers choose which titles to publish, and libraries choose which titles to add to their catalogs. You might not agree with their reasoning or results, but you do have to admit that there are human beings at the helm.
(By the way, Google's library program was put on hold because of criticism. This isn't because it's a bad idea or anything, but rather because the traditional publishing industries got scared they'd lose money. It's an I-was-here-first money-by-copyright battle.)
For many, this filtering-by-design is a major disappointment, which explains the astounding popularity of the World Wide Web. The Web allows everyone to speak up: to post pictures of their pets and babies, their ideas about government, their Harry Potter fan fiction. But in a room where everyone is shouting, nothing gets heard. Northern Light and Yahoo earned their money offering a way through the noise. Other companies, calling themselves search optimization experts, profit by offering their clients the means to be noticed by these search engines, the hypertext equivalent of megaphones.
The filtering process is missing.
Okay, yes, adding content to search databases is a good thing. I might make fun of sport fanaticism, but the desire to retrieve information of choice is a valuable privilege of the individual. I might not care about the Red Sox, but I respect that there are others who do. I also feel extremely happy for the researchers who now have access to volumes of scientific research. In many ways, adding content to a database is like translating content into new languages. Really, these are not bad things.
Even so, we are only adding to the number of people shouting in a room. Google and Yahoo are definitely improving the scope of what we can find, but they are not improving our ability to find. I might proudly accumulate more and more in my attic, while simultaneously making it harder and harder to retrieve anything.
Here’s a real example. My wife and I had been expecting our first child (15 months ago). We were struggling in our decision of her middle name. When we searched for “baby names” on the Web, with quotation marks around the phrase, we found 2.5 million sites at Google, 0.9 million at Yahoo, 0.7 million at MSN. Now, imagine that all the contents of library books and scientific articles and sports broadcasters are added to the Web. Although there may exist a few anthropology articles that would have helped us choose a name, I sincerely believe that over 99% of this new content would have proven unhelpful. I also believe that some of this unhelpful content includes the unusual phrase “baby names.” For example, consider this sentence, which appeared on the Web in November 2004: “Julia Roberts now joins the list of celebrities who have jumped on the Hollywood bandwagon, which gives license to choosing odd baby names.”
After my wife and I decide on a candidate name, we search for that name online, looking for its meaning. The query “Ryan meaning” (without quotes) for the boy’s name Ryan gets 1.2 million hits at Google. (By the way, Google suppresses near-duplicated content and never displays beyond the first 1000 results, so the “true” result set is quite inaccessible.) Because Ryan is a common name, it likely appears numerous times within bibliographies. The word meaning is also extremely common among scholarly articles. If Google indeed adds university libraries to it’s already large database, 1.2 million will become a very small number. As the scientists benefit from a library search, my wife and I will find it that much harder to learn about a particular name using Google.
It is common knowledge among library scientists and search engine experts that you cannot improve the accuracy of a search at the same time you improve its comprehensiveness. Either you get perfect relevance but miss something useful, or you get everything you want along with content that you don’t. As search engines trend toward larger and larger databases, results pages grow more cluttered.
Please, Sir, May I Have Some Less?
Google’s popularity as a search engine has nothing to do with the numbers of results. When I ask people why they like Google (or whatever search engine they prefer), they answer, “Because what I want is usually within the first few results.” I’ve also gotten the answer, “Because it thinks the way that I do.” Most people don’t want millions of results. They want three. Three quality results.
I can’t remember the last time I heard about a search engine improving its algorithms. Perhaps they do this all the time in secret, inventing features behind the scenes. I do know that if a search engine started regularly serving up nonsense, it would go out of business.
So why are these efforts at improving quality so unpronounced? Did you know that while Google pays attention to quotation marks, Lycos doesn’t? That you can type a zip code into Google to get a map? Many of Google’s best features are published in books like Google Hacks and even Google Maps Hacks (O’Reilly & Associates, 2004 and 2006 respectively), where few people are going to look for them. Either nobody cares, or nobody knows the difference.
But when Google adds sports commentary to its search engine, watch out! The story appears in all the major newspapers.
At times like these, I get rather discouraged. I feel as though I am trying to hold back the ocean. It takes me more than a week to index a single, average book. The information world is growing at such an insane pace, my job seems absurd.
At times like these, I have to remind myself of two perspective. First, context. When I write the index for a book with 350 pages, it doesn’t matter that the “book of Google” has over 8 billion. Someone decided that these 350 pages needed to be written, and it’s my job to make them accessible. My work improves this book. No, I haven’t changed the world, but I have made a difference within the context of this one book, in a segment of this one industry, to a small set of readers. For me, indexing is like the civic duty of voting: few win by one vote, and yet every vote counts. It’s also contagious, because voting begets voting. And indexing does beget indexing, because 5% of the people I talk to about my job want to know more, offer me work, or express a desire to become an indexer themselves.
The second perspective is one of application. I don’t have to index books. If I wanted to make a difference at the source, there are many other applications of skills.
The key to these perspectives is a willingness to become activists. We are environmentalists in an information world. Just as scientists show concern with over a one-degree rise in ocean temperature, so should we show concern with a one-percent increase in information dissemination. Bulk up our search engines? This is not an environmentally friendly choice.
I want a search engine—it doesn’t even have to be Google—to announce that they’ve found a way to help me filter out the pages I couldn’t possibly want. The important word here is announce. The modus operandi of these companies is to bring more shouting people into the room, and then publicize this with pride. No wonder the libraries and publishers are in trouble: they’re not being praised for what they do. Neither are the indexers. In the public media, quality gets a whole lot less attention than quantity.
If the search engine is being improved, they’re not telling anyone. Apparently it’s a secret. I don’t want to hear that someone has added billions of pages to the database, unless I hear also about a system that filters billions of pages away.
I think it’s wonderful that more esoteric content is being added to the database. I applaud search engine companies who continue to improve their algorithms. What drives me crazy is that everyone is talking about the first, but not the second. The publicity is lopsided. Why won’t anyone talk about quality any more?
When we asked search engine companies for more, that’s what we got. And we lost precision. Maybe it’s time for us to ask for less.
Press releases from Google, like “Google Checks Out Library Books” [December 14, 2004] and “Google Tunes Into TV” [January 25, 2005], hit the media waves in a grand style. For the first time, entire libraries of books, from Harvard and Stanford Universities to the Universities of Michigan and Oxford, and soon the New York City Public Library, will be available from Google’s website. Within limits of copyright, the words of entire books can be searched. Information philosophers are all over this story, essentially declaring that Google will become the public library of the next generation, excited about how the very nature of libraries might change, and scratching their heads over how the book publishing industry is going to survive yet another hit in the market.
In the second, Google (as well as Yahoo!) applauds themselves for once again providing access to a greater diversity of the world’s information, because television’s closed captioning content has been indexed into a Google Video database. Viewers of public broadcasting and basketball are early adopters, and why not? Finally, all those oh-so-deprived sports consumers can satisfy themselves on more than just the videos, statistics databases, press releases, articles, blogs, commentaries, and (don’t forget) live games themselves. Because now they can search among the announcers’ words.
Now when I search for 76ers, instead of getting 1.61 million hits, I’ll get 1.62 million. Phooey.
Our instinctive reaction is to be impressed. I’m thinking of all those Ph.D. theses gathering dust in the Physics-Optics-Astronomy Library at my alma mater, 150-page books without indexes. I’m thinking about Red Sox fans who, for the first time in a very long time, are interested in the World Series.
But I’m also thinking about the catalog search system at my public library, which won’t improve with Google’s additions. For a book already in the catalog, adding its content doesn’t help at all. Instead, we’d be cluttering up the database with a few trillion new words.
So our instincts are wrong.
Why Quantity Hurts
Every few years, search engine companies are finding new ways to promote themselves by bragging about how much they can find. In the early 1990s, Northern Light was independently rated top among competitors because they searched the largest percentage of the World Wide Web: sixteen percent. Time and Newsweek contributors warned, “We’re not finding everything.”
In the late 1990s, articles about the “invisible web” appeared in popular magazines and newspapers, explaining how search engines cataloged only text, image, and sound files, thus skipping over the good content stored as spreadsheets, databases, and fonts. Again came the cry, “We’re not finding everything!”
And now, in just two months, Google and Yahoo have added more to the huge pile of information: library books and closed captioning data. No longer are our searches limited to the billions of files already on the Web. Yay!
It’s all about quantity. Nobody seems to care about quality any more.
Libraries have been struggling to redefine themselves ever since the Internet (and more, the World Wide Web) reached people’s homes. Book publishers also have suffered. The failure isn’t that of these institutions and industries, however, but of the public. The public seems unaware of the natural filtering process inherent in human behavior. Publishers choose which titles to publish, and libraries choose which titles to add to their catalogs. You might not agree with their reasoning or results, but you do have to admit that there are human beings at the helm.
(By the way, Google's library program was put on hold because of criticism. This isn't because it's a bad idea or anything, but rather because the traditional publishing industries got scared they'd lose money. It's an I-was-here-first money-by-copyright battle.)
For many, this filtering-by-design is a major disappointment, which explains the astounding popularity of the World Wide Web. The Web allows everyone to speak up: to post pictures of their pets and babies, their ideas about government, their Harry Potter fan fiction. But in a room where everyone is shouting, nothing gets heard. Northern Light and Yahoo earned their money offering a way through the noise. Other companies, calling themselves search optimization experts, profit by offering their clients the means to be noticed by these search engines, the hypertext equivalent of megaphones.
The filtering process is missing.
Okay, yes, adding content to search databases is a good thing. I might make fun of sport fanaticism, but the desire to retrieve information of choice is a valuable privilege of the individual. I might not care about the Red Sox, but I respect that there are others who do. I also feel extremely happy for the researchers who now have access to volumes of scientific research. In many ways, adding content to a database is like translating content into new languages. Really, these are not bad things.
Even so, we are only adding to the number of people shouting in a room. Google and Yahoo are definitely improving the scope of what we can find, but they are not improving our ability to find. I might proudly accumulate more and more in my attic, while simultaneously making it harder and harder to retrieve anything.
Here’s a real example. My wife and I had been expecting our first child (15 months ago). We were struggling in our decision of her middle name. When we searched for “baby names” on the Web, with quotation marks around the phrase, we found 2.5 million sites at Google, 0.9 million at Yahoo, 0.7 million at MSN. Now, imagine that all the contents of library books and scientific articles and sports broadcasters are added to the Web. Although there may exist a few anthropology articles that would have helped us choose a name, I sincerely believe that over 99% of this new content would have proven unhelpful. I also believe that some of this unhelpful content includes the unusual phrase “baby names.” For example, consider this sentence, which appeared on the Web in November 2004: “Julia Roberts now joins the list of celebrities who have jumped on the Hollywood bandwagon, which gives license to choosing odd baby names.”
After my wife and I decide on a candidate name, we search for that name online, looking for its meaning. The query “Ryan meaning” (without quotes) for the boy’s name Ryan gets 1.2 million hits at Google. (By the way, Google suppresses near-duplicated content and never displays beyond the first 1000 results, so the “true” result set is quite inaccessible.) Because Ryan is a common name, it likely appears numerous times within bibliographies. The word meaning is also extremely common among scholarly articles. If Google indeed adds university libraries to it’s already large database, 1.2 million will become a very small number. As the scientists benefit from a library search, my wife and I will find it that much harder to learn about a particular name using Google.
It is common knowledge among library scientists and search engine experts that you cannot improve the accuracy of a search at the same time you improve its comprehensiveness. Either you get perfect relevance but miss something useful, or you get everything you want along with content that you don’t. As search engines trend toward larger and larger databases, results pages grow more cluttered.
Please, Sir, May I Have Some Less?
Google’s popularity as a search engine has nothing to do with the numbers of results. When I ask people why they like Google (or whatever search engine they prefer), they answer, “Because what I want is usually within the first few results.” I’ve also gotten the answer, “Because it thinks the way that I do.” Most people don’t want millions of results. They want three. Three quality results.
I can’t remember the last time I heard about a search engine improving its algorithms. Perhaps they do this all the time in secret, inventing features behind the scenes. I do know that if a search engine started regularly serving up nonsense, it would go out of business.
So why are these efforts at improving quality so unpronounced? Did you know that while Google pays attention to quotation marks, Lycos doesn’t? That you can type a zip code into Google to get a map? Many of Google’s best features are published in books like Google Hacks and even Google Maps Hacks (O’Reilly & Associates, 2004 and 2006 respectively), where few people are going to look for them. Either nobody cares, or nobody knows the difference.
But when Google adds sports commentary to its search engine, watch out! The story appears in all the major newspapers.
At times like these, I get rather discouraged. I feel as though I am trying to hold back the ocean. It takes me more than a week to index a single, average book. The information world is growing at such an insane pace, my job seems absurd.
At times like these, I have to remind myself of two perspective. First, context. When I write the index for a book with 350 pages, it doesn’t matter that the “book of Google” has over 8 billion. Someone decided that these 350 pages needed to be written, and it’s my job to make them accessible. My work improves this book. No, I haven’t changed the world, but I have made a difference within the context of this one book, in a segment of this one industry, to a small set of readers. For me, indexing is like the civic duty of voting: few win by one vote, and yet every vote counts. It’s also contagious, because voting begets voting. And indexing does beget indexing, because 5% of the people I talk to about my job want to know more, offer me work, or express a desire to become an indexer themselves.
The second perspective is one of application. I don’t have to index books. If I wanted to make a difference at the source, there are many other applications of skills.
The key to these perspectives is a willingness to become activists. We are environmentalists in an information world. Just as scientists show concern with over a one-degree rise in ocean temperature, so should we show concern with a one-percent increase in information dissemination. Bulk up our search engines? This is not an environmentally friendly choice.
I want a search engine—it doesn’t even have to be Google—to announce that they’ve found a way to help me filter out the pages I couldn’t possibly want. The important word here is announce. The modus operandi of these companies is to bring more shouting people into the room, and then publicize this with pride. No wonder the libraries and publishers are in trouble: they’re not being praised for what they do. Neither are the indexers. In the public media, quality gets a whole lot less attention than quantity.
If the search engine is being improved, they’re not telling anyone. Apparently it’s a secret. I don’t want to hear that someone has added billions of pages to the database, unless I hear also about a system that filters billions of pages away.
I think it’s wonderful that more esoteric content is being added to the database. I applaud search engine companies who continue to improve their algorithms. What drives me crazy is that everyone is talking about the first, but not the second. The publicity is lopsided. Why won’t anyone talk about quality any more?
When we asked search engine companies for more, that’s what we got. And we lost precision. Maybe it’s time for us to ask for less.
Labels: books, Google, power of information, search engines
24 March 2006
The granularity of an online "page number"
When writing a hyperlinked index (where hyperlinks are used instead of page numbers), to what should those links point?
Some people think they should point to the section title in which the information is provided; other people like to point right to the specific word used in the index. The "answer," obviously, is that hyperlinked index entries should take the readers to where the information is, right? The problem -- the reason there's this question of "where do I point my entries" in the first place is that readers of hypertext might find themselves bounced somewhere they don't understand. How many times have you followed a link, only to find yourself fiddling around with the scroll bar to figure out where you ended up? Following a hyperlink is like being blindfolded and transported to an unknown destination.
What you may not know is that book indexes aren't much different. :-)
Think about how book indexes actually work, and you realize that direct readers to the page on which the information starts. An entry like "buoyancy, 164" tells the reader to look somewhere on page 164; an entry like "global harmony, 164-167" tells the reader to start looking somewhere on page 164. The granularity of an index is defined as the smallest unit of area that can pointed to. For printed indexes, this area is the page number. Rarely will you find locators that use fractional or qualified page numbers like 164-1/2 or 164top. (There are such things as qualified locators, like 164f, which might point to the footnote on page 164, but even in the books in which they're used they comprise only a small number of all locators used.)
If you follow the standards of the industry, then, the granularity of a printed index is one physical page. For this reason, books that have lots of words on a page -- big pages, narrow margins, tiny print -- are less friendly to book indexers. It's like telling someone that there's a needle in that 164th haystack over there. Maybe we should count our blessings that someone bothered to number the haystacks, but ideally this is where the book designer starts earning her salary. Book pages don't have to look like haystacks -- more accurately, wordstacks -- if the book has legible headings and subheadings. Books can be written with quickly visible landmarks within the pages, like italics and boldface, larger and smaller font sizes, headings and callouts, footnotes, and so on. Going back to the blindfolded analogy, there's no reason we have to drop our readers into deserts of information, when we can drop them in a place surrounded by location clues and navigational signs, like at a train station.
On the Web, however, there is no such thing as a printed page. Web pages can be any length, from tiny pop-up windows with only a sentence fragment of information within, to long scrolls of endless paragraphs and images. Additionally, you don't have to direct the reader to just the page any more, but rather you can deposit him anywhere within the page. The granularity of a Web page is a word! You can send someone into the middle of a paragraph.
When you have tiny little windows of information, using that window as a destination is a no-brainer: the reader arrives at a single sentence of information, which is what he needs. It doesn't matter if you point him to the beginning, middle, or end of that sentence, because it's all they get to read. Pointing someone to an isolated window of information -- what Web authors call "chunks" -- is as easy as looking into a food pantry that contains only a single can. But when you have longer pages, and you have the ability to point someone to any spot within those longer pages, you have a decision to make. And it's a decision that didn't exist in the printed world, with its larger granularity.
The solution is to connect the text of the index entry with the text of the documentation. Not the meaning, but the actual words. If the index entries are written to almost identically match those of the documentation, then the reader won't mind as much because it won't look like a desert. They'll have exactly the landmark they need right in front of them. The entry "cancer, prevention of," for example, could point directly to this line without a problem:
... cessation of smoking. In fact, many physicians are well aware that one way to prevent cancer is to quit ...
That's because the words of your index entry, which are cancer and prevention, appear almost verbatim in that line of text. And if this information were part of a section titled "Using Peer Pressure to Help Patients Quit Smoking," then you really wouldn't want to point to the heading for context. That's because it's unclear to the reader that you're actually directing him to information about cancer or prevention. You're making them work at it.
And then there's the other situation. Using the same sentence and heading as above, where should the indexer point readers who look up the entry "smoking, how to quit"? Clearly they should go right to the heading. If they went to the line that talked about physicians, they wouldn't know where they are.
Our original question here was this: When writing a hyperlinked index (where hyperlinks are used instead of page numbers), to what should those links point? Clearly the only way to answer this question comprehensively is to suggest that the language of hyperlink indexes has two contexts: the index entry itself and the destination location. These two contexts need to work together. And as we saw, the same is true with the printed book: having arrived at page 164, how quickly can you find the idea you were looking for?
Looking this closely at hyperlinked indexes only emphasizes something we need for all indexing: use index entries that match the documentation text. If you have to write a slightly longer entry, that's okay. Instead of "cigarettes," use "cigarette smoking, quitting." Instead of "social networks," use "social networks and peer pressure." The people who work with search engines and Internet marketing are familiar with the term trigger words, which refers to visible language that matches the mental language of the searcher. If you're thinking of the words "white elephant," then a result of "pale pachyderm" doesn't work because it doesn't trigger your sense of recognition.
So the next time there's a white elephant in haystack 164, be sure to tell someone as explicitly as possible.
Some people think they should point to the section title in which the information is provided; other people like to point right to the specific word used in the index. The "answer," obviously, is that hyperlinked index entries should take the readers to where the information is, right? The problem -- the reason there's this question of "where do I point my entries" in the first place is that readers of hypertext might find themselves bounced somewhere they don't understand. How many times have you followed a link, only to find yourself fiddling around with the scroll bar to figure out where you ended up? Following a hyperlink is like being blindfolded and transported to an unknown destination.
What you may not know is that book indexes aren't much different. :-)
Think about how book indexes actually work, and you realize that direct readers to the page on which the information starts. An entry like "buoyancy, 164" tells the reader to look somewhere on page 164; an entry like "global harmony, 164-167" tells the reader to start looking somewhere on page 164. The granularity of an index is defined as the smallest unit of area that can pointed to. For printed indexes, this area is the page number. Rarely will you find locators that use fractional or qualified page numbers like 164-1/2 or 164top. (There are such things as qualified locators, like 164f, which might point to the footnote on page 164, but even in the books in which they're used they comprise only a small number of all locators used.)
If you follow the standards of the industry, then, the granularity of a printed index is one physical page. For this reason, books that have lots of words on a page -- big pages, narrow margins, tiny print -- are less friendly to book indexers. It's like telling someone that there's a needle in that 164th haystack over there. Maybe we should count our blessings that someone bothered to number the haystacks, but ideally this is where the book designer starts earning her salary. Book pages don't have to look like haystacks -- more accurately, wordstacks -- if the book has legible headings and subheadings. Books can be written with quickly visible landmarks within the pages, like italics and boldface, larger and smaller font sizes, headings and callouts, footnotes, and so on. Going back to the blindfolded analogy, there's no reason we have to drop our readers into deserts of information, when we can drop them in a place surrounded by location clues and navigational signs, like at a train station.
On the Web, however, there is no such thing as a printed page. Web pages can be any length, from tiny pop-up windows with only a sentence fragment of information within, to long scrolls of endless paragraphs and images. Additionally, you don't have to direct the reader to just the page any more, but rather you can deposit him anywhere within the page. The granularity of a Web page is a word! You can send someone into the middle of a paragraph.
When you have tiny little windows of information, using that window as a destination is a no-brainer: the reader arrives at a single sentence of information, which is what he needs. It doesn't matter if you point him to the beginning, middle, or end of that sentence, because it's all they get to read. Pointing someone to an isolated window of information -- what Web authors call "chunks" -- is as easy as looking into a food pantry that contains only a single can. But when you have longer pages, and you have the ability to point someone to any spot within those longer pages, you have a decision to make. And it's a decision that didn't exist in the printed world, with its larger granularity.
The solution is to connect the text of the index entry with the text of the documentation. Not the meaning, but the actual words. If the index entries are written to almost identically match those of the documentation, then the reader won't mind as much because it won't look like a desert. They'll have exactly the landmark they need right in front of them. The entry "cancer, prevention of," for example, could point directly to this line without a problem:
... cessation of smoking. In fact, many physicians are well aware that one way to prevent cancer is to quit ...
That's because the words of your index entry, which are cancer and prevention, appear almost verbatim in that line of text. And if this information were part of a section titled "Using Peer Pressure to Help Patients Quit Smoking," then you really wouldn't want to point to the heading for context. That's because it's unclear to the reader that you're actually directing him to information about cancer or prevention. You're making them work at it.
And then there's the other situation. Using the same sentence and heading as above, where should the indexer point readers who look up the entry "smoking, how to quit"? Clearly they should go right to the heading. If they went to the line that talked about physicians, they wouldn't know where they are.
Our original question here was this: When writing a hyperlinked index (where hyperlinks are used instead of page numbers), to what should those links point? Clearly the only way to answer this question comprehensively is to suggest that the language of hyperlink indexes has two contexts: the index entry itself and the destination location. These two contexts need to work together. And as we saw, the same is true with the printed book: having arrived at page 164, how quickly can you find the idea you were looking for?
Looking this closely at hyperlinked indexes only emphasizes something we need for all indexing: use index entries that match the documentation text. If you have to write a slightly longer entry, that's okay. Instead of "cigarettes," use "cigarette smoking, quitting." Instead of "social networks," use "social networks and peer pressure." The people who work with search engines and Internet marketing are familiar with the term trigger words, which refers to visible language that matches the mental language of the searcher. If you're thinking of the words "white elephant," then a result of "pale pachyderm" doesn't work because it doesn't trigger your sense of recognition.
So the next time there's a white elephant in haystack 164, be sure to tell someone as explicitly as possible.
Labels: books, indexing process, keywording, pages and page ranges, web indexing
