Tuesday, December 8, 2009
Monday, November 30, 2009
Muddiest Point
Are there any non-commercial open source search engines that provide an alternative to google?
Reading Notes Dec. 1, 2009
Using a Wiki to Manage a Library Instruction Program
This article discusses a program to integrate wikis into library instruction at the library at East Tennessee State University, where the author is employed as reference librarian -- wikis are used to plan and expand on library instruction sessions -- professors can expand on specific questions they have during information sessions, or can help students to specify assignments for an instruction session.
Also, wikis can be used as a collaborative tool for exchanging information among instruction librarians. In addition, the article includes some helpful links to commercial sites with software allowing you to create your own wiki, as well as some background literature on wikis.
Creating the Academic Library Folksonomy
The article discusses how users can create folksonomy -- taxonomies of keywords created and shared by lots of users. Such folksonomies can help catalog useful internet sites, and integrate these folksonomies into academic libraries. This can provide quality indices of web sites, and also bring "grey" literature to light. The article gives a few examples of libraries that are trying out social tagging, f.e. the University of Pennsylvania, or Stanford University. Services uch as delicio or citeulike allow users to share these tags.
The article also discusses very briefly some of the risks and concerns about social tagging, for example spagging (spam tagging), or the variation of tags and the lack of controlled vocabulary. There are other concerns, not mentioned by the article, for example that tagging might provide institutions with a rationale to cut cataloging jobs, a discussion which has been held at the LC for a while, and an ongoing discussion whether or not it is useful to have controlled vocabularies to describe collections.
Weblogs: Their Use and Application in Science and Technology Libraries
Discusses the history of weblogs beginning in the early 1990s. Web logs are web sites that resemble personal journals, which allow visitors to comment on entries, and also include an archive of past weblogs. The authors highlight that web logs really took off in the late 1990s, when software package became available that allow users to build their own blogs. The article highlights web logs as a collaborative tool in science and technology libraries, which has some advatages over, for example email, since it is more easily searchable and can be organized according to specific subjects. RSS feeds can serve as a reminder to users to visit blogs. The article also explores reference blogs as an alternative to email. In blog instruction, librarians can have an important function, and train students in setting up and maintaining blogs. The article is from 2004 -- since then, many libraries have integrated blogs into their services..
Wikipedia
Interview with Jimmy Wales on Wikipedia -- the question here is to what extent the ideal of wikipedia matches the reality -- there was an article in Times Magazine, which Donna Guerin recommended for our LIS 2000 class, about some of the discrepancies: Wikipedia: A Victim of Its Own Success?
http://www.time.com/time/magazine/article/0,9171,1924492-1,00.html
These discrepancies suggest to take Jimmy Wales claims about wikipedia as this diverse, global, bottom-up project with a grain of salt, and rather reflect about persistent hierarchies in the way knowledge is created and organized.
This article discusses a program to integrate wikis into library instruction at the library at East Tennessee State University, where the author is employed as reference librarian -- wikis are used to plan and expand on library instruction sessions -- professors can expand on specific questions they have during information sessions, or can help students to specify assignments for an instruction session.
Also, wikis can be used as a collaborative tool for exchanging information among instruction librarians. In addition, the article includes some helpful links to commercial sites with software allowing you to create your own wiki, as well as some background literature on wikis.
Creating the Academic Library Folksonomy
The article discusses how users can create folksonomy -- taxonomies of keywords created and shared by lots of users. Such folksonomies can help catalog useful internet sites, and integrate these folksonomies into academic libraries. This can provide quality indices of web sites, and also bring "grey" literature to light. The article gives a few examples of libraries that are trying out social tagging, f.e. the University of Pennsylvania, or Stanford University. Services uch as delicio or citeulike allow users to share these tags.
The article also discusses very briefly some of the risks and concerns about social tagging, for example spagging (spam tagging), or the variation of tags and the lack of controlled vocabulary. There are other concerns, not mentioned by the article, for example that tagging might provide institutions with a rationale to cut cataloging jobs, a discussion which has been held at the LC for a while, and an ongoing discussion whether or not it is useful to have controlled vocabularies to describe collections.
Weblogs: Their Use and Application in Science and Technology Libraries
Discusses the history of weblogs beginning in the early 1990s. Web logs are web sites that resemble personal journals, which allow visitors to comment on entries, and also include an archive of past weblogs. The authors highlight that web logs really took off in the late 1990s, when software package became available that allow users to build their own blogs. The article highlights web logs as a collaborative tool in science and technology libraries, which has some advatages over, for example email, since it is more easily searchable and can be organized according to specific subjects. RSS feeds can serve as a reminder to users to visit blogs. The article also explores reference blogs as an alternative to email. In blog instruction, librarians can have an important function, and train students in setting up and maintaining blogs. The article is from 2004 -- since then, many libraries have integrated blogs into their services..
Wikipedia
Interview with Jimmy Wales on Wikipedia -- the question here is to what extent the ideal of wikipedia matches the reality -- there was an article in Times Magazine, which Donna Guerin recommended for our LIS 2000 class, about some of the discrepancies: Wikipedia: A Victim of Its Own Success?
http://www.time.com/time/magazine/article/0,9171,1924492-1,00.html
These discrepancies suggest to take Jimmy Wales claims about wikipedia as this diverse, global, bottom-up project with a grain of salt, and rather reflect about persistent hierarchies in the way knowledge is created and organized.
Monday, November 23, 2009
Monday, November 16, 2009
Sunday, November 15, 2009
Reading Notes November 17
Shreeves, S. L., Habing, T. O., Hagedorn, K., & Young, J. A. (2005). Current developments and future trends for the OAI protocol for metadata harvesting. Library Trends, 53(4), 576-589.
This text was not easy to understand, but here are some excerpts:
"The mission of the Open Archives Initiative (...) is to "develop and promote interoperability standards that aim to facilitate the efficient dissemination of content" (Open Archives Initiative, n.d. a). The Protocol for Metadata Harvesting, a tool developed through the OAI, facilitates interoperability between disparate and diverse collections of metadata through a relatively simple protocol based on common standards (XML, HTTP, and Dublin Core)."
"The OAI protocol requires that data providers expose metadata in at least unqualified Dublin Core; however, the use of other metadata schmas is possible and encouraged. The protocol can provide access to parts of the "invisible Web" that are not easily accessible to search engines
(such as resources within databases) (Sherman & Price, 2003) and can provide ways for communities of interest to aggregate resources from geographically diffuse collections."
-- The data is searchable and browsable without any manual cataloging of the various OAI repositories.
-- I could not pull up the text on how search engines work today, and will continue the reading tomorrow, once I get access to another computer.
White Paper: The Deep Web: Surfacing Hidden Value
Michael K. Bergman
Journal of Electronic Publishing, vol. 7, no. 1, August, 2001
DOI: http://dx.doi.org/10.3998/3336451.0007.104
-- This article is very enlightening -- it covers a technology (DeepPlanet, which can search the deep web)
-Most of the Web's information is buried far down on dynamically generated sites, and standard search engines never find it-- it si about 500 times the size of the surface web
--Traditional search engines create their indices by spidering or crawling surface Web pages. To be discovered, the page must be static and linked to other pages. Traditional search engines can not "see" or retrieve content in the deep Web
--The deep Web is qualitatively different from the surface Web. Deep Web sources store their content in searchable databases that only produce results dynamically in response to a direct request.
-- The deep web is a very coveted commodity
-- Study by NEC research initiative that search engine sby google and northern light only crawl 16% of the web's content
--Search engines obtain their listings in two ways: Authors may submit their own Web pages, or the search engines "crawl" or "spider" documents by following one hypertext link to another. The latter returns the bulk of the listings. Crawlers work by recording every hypertext link in every page they index crawling.
-- The crawls used to be indiscriminate, but "the most recent generation of search engines (notably Google) have replaced the random link-following approach with directed crawling and indexing based on the "popularity" of pages. In this approach, documents more frequently cross-referenced than other documents are given priority both for crawling and in the presentation of results. This approach provides superior results when simple queries are issued, but exacerbates the tendency to overlook documents with few links."
__ the problem here: without a linkage from another Web document, a page will never be discovered.
-- They don't use the term invisible web -- it is not invisible, but rather unindexable
-- the article continues to describe the study in more detail
--it is impossible to completely index the deep content, but new technologies need to be developed to search the complete web
Searching must evolve to encompass the complete Web.
This text was not easy to understand, but here are some excerpts:
"The mission of the Open Archives Initiative (...) is to "develop and promote interoperability standards that aim to facilitate the efficient dissemination of content" (Open Archives Initiative, n.d. a). The Protocol for Metadata Harvesting, a tool developed through the OAI, facilitates interoperability between disparate and diverse collections of metadata through a relatively simple protocol based on common standards (XML, HTTP, and Dublin Core)."
"The OAI protocol requires that data providers expose metadata in at least unqualified Dublin Core; however, the use of other metadata schmas is possible and encouraged. The protocol can provide access to parts of the "invisible Web" that are not easily accessible to search engines
(such as resources within databases) (Sherman & Price, 2003) and can provide ways for communities of interest to aggregate resources from geographically diffuse collections."
-- The data is searchable and browsable without any manual cataloging of the various OAI repositories.
-- I could not pull up the text on how search engines work today, and will continue the reading tomorrow, once I get access to another computer.
White Paper: The Deep Web: Surfacing Hidden Value
Michael K. Bergman
Journal of Electronic Publishing, vol. 7, no. 1, August, 2001
DOI: http://dx.doi.org/10.3998/3336451.0007.104
-- This article is very enlightening -- it covers a technology (DeepPlanet, which can search the deep web)
-Most of the Web's information is buried far down on dynamically generated sites, and standard search engines never find it-- it si about 500 times the size of the surface web
--Traditional search engines create their indices by spidering or crawling surface Web pages. To be discovered, the page must be static and linked to other pages. Traditional search engines can not "see" or retrieve content in the deep Web
--The deep Web is qualitatively different from the surface Web. Deep Web sources store their content in searchable databases that only produce results dynamically in response to a direct request.
-- The deep web is a very coveted commodity
-- Study by NEC research initiative that search engine sby google and northern light only crawl 16% of the web's content
--Search engines obtain their listings in two ways: Authors may submit their own Web pages, or the search engines "crawl" or "spider" documents by following one hypertext link to another. The latter returns the bulk of the listings. Crawlers work by recording every hypertext link in every page they index crawling.
-- The crawls used to be indiscriminate, but "the most recent generation of search engines (notably Google) have replaced the random link-following approach with directed crawling and indexing based on the "popularity" of pages. In this approach, documents more frequently cross-referenced than other documents are given priority both for crawling and in the presentation of results. This approach provides superior results when simple queries are issued, but exacerbates the tendency to overlook documents with few links."
__ the problem here: without a linkage from another Web document, a page will never be discovered.
-- They don't use the term invisible web -- it is not invisible, but rather unindexable
-- the article continues to describe the study in more detail
--it is impossible to completely index the deep content, but new technologies need to be developed to search the complete web
Searching must evolve to encompass the complete Web.
Subscribe to:
Posts (Atom)