Exploring the World Digital Library
1. http://www.flickr.com/photos/42354457@N07/3998337729/?addedcomment=1#comment72157622557339392
2. http://www.flickr.com/photos/42354457@N07/3999108966/sizes/o/
3. http://www.flickr.com/photos/42354457@N07/3999121772/
4. http://www.flickr.com/photos/42354457@N07/3999129796/
5. http://www.flickr.com/photos/42354457@N07/3998371253/
6. http://www.flickr.com/photos/42354457@N07/3999149030/
7. http://www.flickr.com/photos/42354457@N07/3999207116/
Link to my thriller on the World Digital Library:
http://www.screencast.com/users/Khering145/folders/Jing/media/867a7e7d-2926-49ac-8306-a3a4f0b128fe
Saturday, October 10, 2009
Saturday, October 3, 2009
Assignment 3: CiteUlike and Zotero
My citeulike url:
http://www.citeulike.org/user/khering/
The import from zotero/google books is called zotero, or file-import-09-10-03 (generic) and the one from citeulike is called citeulike_import
My three main collections are digital-preservation; audio-preservation; curating-oral-history plus additional tags.
http://www.citeulike.org/user/khering/
The import from zotero/google books is called zotero, or file-import-09-10-03 (generic) and the one from citeulike is called citeulike_import
My three main collections are digital-preservation; audio-preservation; curating-oral-history plus additional tags.
Friday, October 2, 2009
Week 5/6 (Sept. 29-Oct. 6)
Muddiest Point:
This is an issue regarding compression and preservation of different file types which I have encountered when working with multimedia files in an archive. Is it possible for a computer to detect previous versions of a file that was saved as a different file type? For example, if you download an image as a .jpg file, import it into photoshop, alter it, and then save it as a .gif file or .tiff file, can a computer theoretically detect that the .gif file used to be a .jpg file? Such a trail is important for preservation, because the extension might obscure the fact that a file might have been compressed in a lossy format before, so at first sight it might look as if a file is compressed in a lossless format, even though it has been compressed in a lossy format before...
Reading Notes:
This week was the week of networks, both in LIS 2000 and in this class, and I am feeling a bit networked-out. Like many people, I have become acquainted with LAN networks at work while crawling under tables to disentangle some amorphous cable masses to figure out why a printer got stuck. As far as I understand it, LAN networks still depend on hard wires (co-axial cables as we learned). What is important here is that LAN networks do not depend on leased telecommunication networks, which gives the owners more control. As far as I understood, ethernet is a technology that enables LAN (the history of ethernet was actually fairly interesting, too -- I was not aware it was invented by R. Metcalfe at XEROX.) The article on the variety of computer networks was dizzying -- I was particularly interested in MAN networks -- who has control over MAN's -- cities or towns or private companies? I didn't know that the internet is short for internetwork. I want to know more about the physical infrastructure of the internet -- where are the hubs located? Important is the mix of private and public networks, which politicizes the whole issue -- who controls the access to these networks? And how are libraries connected to them? Does the U of Pittsburgh have a CAN, by the way?
RFID
this was a very informative article with a pragmatic perspective on RFID -- as a technology, it brings advantages, but also new pressures to increase efficiency (and potential job cuts) -- I also found her reflections on the rationale of libraries to introduce new technology very enlightening: a technology becomes introduced, is around, and libraries have to adapt and deal with these changes -- with RFID it sounds as if the development is going in this direction, if libraries like it or not....
Commented on:
Kristine Harveaux-Lundeen.s blog and rsj2600's blog
This is an issue regarding compression and preservation of different file types which I have encountered when working with multimedia files in an archive. Is it possible for a computer to detect previous versions of a file that was saved as a different file type? For example, if you download an image as a .jpg file, import it into photoshop, alter it, and then save it as a .gif file or .tiff file, can a computer theoretically detect that the .gif file used to be a .jpg file? Such a trail is important for preservation, because the extension might obscure the fact that a file might have been compressed in a lossy format before, so at first sight it might look as if a file is compressed in a lossless format, even though it has been compressed in a lossy format before...
Reading Notes:
This week was the week of networks, both in LIS 2000 and in this class, and I am feeling a bit networked-out. Like many people, I have become acquainted with LAN networks at work while crawling under tables to disentangle some amorphous cable masses to figure out why a printer got stuck. As far as I understand it, LAN networks still depend on hard wires (co-axial cables as we learned). What is important here is that LAN networks do not depend on leased telecommunication networks, which gives the owners more control. As far as I understood, ethernet is a technology that enables LAN (the history of ethernet was actually fairly interesting, too -- I was not aware it was invented by R. Metcalfe at XEROX.) The article on the variety of computer networks was dizzying -- I was particularly interested in MAN networks -- who has control over MAN's -- cities or towns or private companies? I didn't know that the internet is short for internetwork. I want to know more about the physical infrastructure of the internet -- where are the hubs located? Important is the mix of private and public networks, which politicizes the whole issue -- who controls the access to these networks? And how are libraries connected to them? Does the U of Pittsburgh have a CAN, by the way?
RFID
this was a very informative article with a pragmatic perspective on RFID -- as a technology, it brings advantages, but also new pressures to increase efficiency (and potential job cuts) -- I also found her reflections on the rationale of libraries to introduce new technology very enlightening: a technology becomes introduced, is around, and libraries have to adapt and deal with these changes -- with RFID it sounds as if the development is going in this direction, if libraries like it or not....
Commented on:
Kristine Harveaux-Lundeen.s blog and rsj2600's blog
Friday, September 25, 2009
Week 4 (or 5) Sept. 22-29
Muddiest Point
I would like to learn more about the underlying structure of various databases I use as a researcher. How does a keyword search in a database look like from a technical perspective? Do keyword searches differ -- from a technical perspective-- from searches based on controlled subject headings, such as the LC subject headings? And how do databases, such as JSTOR, rank results -- what is the basis for the ranking?
Reading notes:
I enjoyed reading the behind-the-scenes report on the production of the digital Imagining Pittsburgh Collection Pittsburgh collection, produced under the lead of the DRL with an IMLS grant. It was not only interesting from a technical perspective, but was also a refreshingly candid project description (often, project reports are rather self-congratulatory and don’t mention the difficulties posed by larger digitizing projects to collaborate, and to deal with technical challenges, content, and different organizational backgrounds all at once). It highlighted the challenges of the three institutions that were collaborating on the project, of agreeing on shared standards, while also serving the individual interests of each institution. A major challenge for many digitization projects is the selection of the images that should be digitized and Galloway underlined that the subject headings that the project created were key for the selection of which images to digitize. After last week’s reading’s, the paragraphs on metadata were interesting, and highlighted how the Dublin Core elements were critical in ensuring the interoperability of the metadata of the individual institutions. From a variety of options, they project participants agreed on using the LC subject headings for the description. Galloway also addressed different workflow challenges, and the difficulties of working with different databases in different projects. They agreed on the quality for the production masters (600 dpi) that ensured the uniform quality of the images. (The quality of the production master also allows to look at different sizes of the image, and to magnify parts of individual images when exploring the collection online). Finally, it outlined the challenges allowing users to find different ways to explore the collections as a whole, and individual images.
I also looked at the site, and the reader can do subject searches, keyword searches, searches by collection. You can explore by time, location, collection, or theme. You can also look at images with captions, with full record, or just captions. It really offers a lot of ways to search and explore.
Has anyone looked the experimental visualization prototype, the Bungee View, in more detail?
Compression
Compression is a huge issue in multi-media collections, so the articles were very enlightening – I didn’t understand all the details about the different algorithms, but it clarified the principles of compression, and, in the section on video compression, the differences between a video file and a video stream. Unfortunately, the link to the part of the article on lossy compression did not work...The advantages of compression are clear – they save space on expensive storage devices. On the other hand, it also creates huge problems for archives, which have to deal with files in x many formats, many of which are in compressed, often proprietary formats, so they aren’t archival quality to begin with. The pressure to compress video files is even greater than for audio files, because they are so big – uncompressed video would take up an enormous server space. So, many archives just don’t have the money to buy all that server space, and have no choice but to save the files in a compressed format. So, in a different way than for paper, space continues to be a huge problem.
Just wanted to double check – once a file is in a compressed, lossy, format, you can not just uncompress the file – the missing data is gone, is it?
Comments
Commented on Tiffany J. Brand's blog:
http://tiffanybrandlis2600.blogspot.com/
And Letisha Goerner's blog:
http://letishagoerner2600.blogspot.com/
I would like to learn more about the underlying structure of various databases I use as a researcher. How does a keyword search in a database look like from a technical perspective? Do keyword searches differ -- from a technical perspective-- from searches based on controlled subject headings, such as the LC subject headings? And how do databases, such as JSTOR, rank results -- what is the basis for the ranking?
Reading notes:
I enjoyed reading the behind-the-scenes report on the production of the digital Imagining Pittsburgh Collection Pittsburgh collection, produced under the lead of the DRL with an IMLS grant. It was not only interesting from a technical perspective, but was also a refreshingly candid project description (often, project reports are rather self-congratulatory and don’t mention the difficulties posed by larger digitizing projects to collaborate, and to deal with technical challenges, content, and different organizational backgrounds all at once). It highlighted the challenges of the three institutions that were collaborating on the project, of agreeing on shared standards, while also serving the individual interests of each institution. A major challenge for many digitization projects is the selection of the images that should be digitized and Galloway underlined that the subject headings that the project created were key for the selection of which images to digitize. After last week’s reading’s, the paragraphs on metadata were interesting, and highlighted how the Dublin Core elements were critical in ensuring the interoperability of the metadata of the individual institutions. From a variety of options, they project participants agreed on using the LC subject headings for the description. Galloway also addressed different workflow challenges, and the difficulties of working with different databases in different projects. They agreed on the quality for the production masters (600 dpi) that ensured the uniform quality of the images. (The quality of the production master also allows to look at different sizes of the image, and to magnify parts of individual images when exploring the collection online). Finally, it outlined the challenges allowing users to find different ways to explore the collections as a whole, and individual images.
I also looked at the site, and the reader can do subject searches, keyword searches, searches by collection. You can explore by time, location, collection, or theme. You can also look at images with captions, with full record, or just captions. It really offers a lot of ways to search and explore.
Has anyone looked the experimental visualization prototype, the Bungee View, in more detail?
Compression
Compression is a huge issue in multi-media collections, so the articles were very enlightening – I didn’t understand all the details about the different algorithms, but it clarified the principles of compression, and, in the section on video compression, the differences between a video file and a video stream. Unfortunately, the link to the part of the article on lossy compression did not work...The advantages of compression are clear – they save space on expensive storage devices. On the other hand, it also creates huge problems for archives, which have to deal with files in x many formats, many of which are in compressed, often proprietary formats, so they aren’t archival quality to begin with. The pressure to compress video files is even greater than for audio files, because they are so big – uncompressed video would take up an enormous server space. So, many archives just don’t have the money to buy all that server space, and have no choice but to save the files in a compressed format. So, in a different way than for paper, space continues to be a huge problem.
Just wanted to double check – once a file is in a compressed, lossy, format, you can not just uncompress the file – the missing data is gone, is it?
Comments
Commented on Tiffany J. Brand's blog:
http://tiffanybrandlis2600.blogspot.com/
And Letisha Goerner's blog:
http://letishagoerner2600.blogspot.com/
Friday, September 18, 2009
Week 3 (or 4), Sept. 15-Sept. 22
Muddiest Point
We touched on this only briefly during the lecture, but I'd like to better understand DSpace and how libraries can use it?
Reading Notes
I was glad I was able to read some of the reading notes from other students -- it helped clarify things a bit -- the Dublin Core Data Model article was bewildering.
My ideas about metadata had been rather foggy, and so the article on metadata was very enlightening. Metadata can reflect content, context, and structure. In libraries, bibliographic metadata provides access to contents, f.e. through indexes and catalogs. In archives and museums, metadata often describes the context of records and enables authentication of records and objects (f.e. accession records). Metadata is not only description, but also relate to administration, accession, and preservation and use of collections. Representing the structure of objects is central to metadata development. Systems should reflect content, but should also include additional information about the content. This also helps to specify the intellectual integrity of objects, and maintaining the relationships between objects. This has become a central aim in the preservation of digital objects. It was particularly interesting to read about the role of metadata in digital preservation, as it ensures that digital information, or information objects as she calls them, will survive through migrations of hardware and software and will be preserved.
It was helpful to get a sense of the different types of databases, but the article was only slightly clearer to me than the Dublin Core Data Model one (I tried reading both from various directions, but it didn't do much good) -- I'd like to see it illustrations/models of various models of databases, maybe that will make things a bit clearer.
Comments
I commented on Annie's LIS 2600 blog:
http://annie-lis2600-at-pitt-blog.blogspot.com/
We touched on this only briefly during the lecture, but I'd like to better understand DSpace and how libraries can use it?
Reading Notes
I was glad I was able to read some of the reading notes from other students -- it helped clarify things a bit -- the Dublin Core Data Model article was bewildering.
My ideas about metadata had been rather foggy, and so the article on metadata was very enlightening. Metadata can reflect content, context, and structure. In libraries, bibliographic metadata provides access to contents, f.e. through indexes and catalogs. In archives and museums, metadata often describes the context of records and enables authentication of records and objects (f.e. accession records). Metadata is not only description, but also relate to administration, accession, and preservation and use of collections. Representing the structure of objects is central to metadata development. Systems should reflect content, but should also include additional information about the content. This also helps to specify the intellectual integrity of objects, and maintaining the relationships between objects. This has become a central aim in the preservation of digital objects. It was particularly interesting to read about the role of metadata in digital preservation, as it ensures that digital information, or information objects as she calls them, will survive through migrations of hardware and software and will be preserved.
It was helpful to get a sense of the different types of databases, but the article was only slightly clearer to me than the Dublin Core Data Model one (I tried reading both from various directions, but it didn't do much good) -- I'd like to see it illustrations/models of various models of databases, maybe that will make things a bit clearer.
Comments
I commented on Annie's LIS 2600 blog:
http://annie-lis2600-at-pitt-blog.blogspot.com/
Thursday, September 10, 2009
Week 2 (Sept. 8-Sept. 15), Assignment 2
Digitizing Images and Uploading them on Flickr
Here is the URL (I think) to my FlickR account. I will add more comments on the digitization process soon.
Flickr URL:
http://www.flickr.com/photos/42354457@N07/sets/72157622333487728/
Notes on the digitization:
I created this mini exhibit on the history of public transportation planning by Pittsburgh citizens in the early 1920s. I don’t have a car, and am often frustrated with public transport options, so I am always interested in learning about people's suggestions for improving public transport in the past. Usually, these suggestions never materialized. So, I found this report in the collections of the university library and scanned it based on the guidelines – master images scanned at 600 dpi. I used greyscale for the text images (much better than B&W photo) and color image for the diagrams. I then used photoshop and reduced the image size twice and created an image for the screen and an image for the thumbnail, and saved both as jpgs. I then uploaded it to Flickr and added tags and comments – it was the first time I created a Flickr account, so it was all new to me. I thought I was the first one to digitize the report, but as it always happens when you think you are the first, you are certainly not. I found out that the report had already been digitized as part of the Historic Pittsburgh Collections. So much for being the pioneer digitizer.
Some notes and problems: What struck me when creating the master image was how big the image is when you create a really good scan (it was exceeding 100 MB’s). This is a pretty big file, and so this is indicating the space problems on your hard drive or storage unit that you encounter if you make really high quality scans. And drive space costs money. So, with limited money and hard drive space, you can definitely not save everything.
I also had some problems with writing the tags for Flickr. Flickr always combined my key words into one big word, so it looks really strange, like: 60yearsbeforethesubwaytherewasaplan. Weird, isn’t it?
Also, I wasn’t sure if it is possible to actually arrange the thumbnails where they belong (at the first page of the exhibit in Flickr). Instead, all my thumbnails are showing up as part of the images that are part of the exhibit and Flickr is creating its own thumbnails. Does anyone have any suggestions?
Reading Notes
I was roughly familiar with the story of Linux, but it was interesting to read the story how it evolved as a spin-off from UNIX, and has been developed as an open source software ever since by a large group of programmers. It was also interesting that Linux programmers became more user friendly over time. I still find it difficult to understand, though. Can you just install it on a windows machine? I browsed through the Mac OSX/Linux article (I suppose we are not supposed to read these technical pieces line by line), and what struck me was Singh's critical, but dispassionate position toward windows as a client computer and his acknowledgement that most people use windows, not the least because it is cheaper than Mac. The update on the windows roadmap reads like a promotion piece – everything is moving forward, steadily improving, if you only trust microsoft!
Overall, the readings about the different operating systems highlight that compatibility is a huge problem – compatibility between older and newer operating systems, and also compatibility between mac and windows – I am not sure about Linux. It also creates a lot of practical problems if you work at a place that has both macs and pcs – at the archives where I used to work we often had to transfer multimedia files back and forth between windows and macs, which was not always easy.
Muddiest point
The lectures were all very clear; I don’t really have a muddy point. I am just generally wondering about the implications of the constant development and updates of hardware for libraries and archives – should libraries keep versions of old hardware models (f.e. 268 PCs with 5 ¼ inch floppy drives) so that they will be able to read files on old floppies that may contain important files and re-emerge one day?
Comments
I am a little unclear on this, but are we supposed to document where we have left comments?
I commented on Kristine Harveux-Lundeen's blog:
http://2600kristineharveaux-lundeen.blogspot.com/
And on Letisha Goerner's blog:
http://letishagoerner2600.blogspot.com/
Here is the URL (I think) to my FlickR account. I will add more comments on the digitization process soon.
Flickr URL:
http://www.flickr.com/photos/42354457@N07/sets/72157622333487728/
Notes on the digitization:
I created this mini exhibit on the history of public transportation planning by Pittsburgh citizens in the early 1920s. I don’t have a car, and am often frustrated with public transport options, so I am always interested in learning about people's suggestions for improving public transport in the past. Usually, these suggestions never materialized. So, I found this report in the collections of the university library and scanned it based on the guidelines – master images scanned at 600 dpi. I used greyscale for the text images (much better than B&W photo) and color image for the diagrams. I then used photoshop and reduced the image size twice and created an image for the screen and an image for the thumbnail, and saved both as jpgs. I then uploaded it to Flickr and added tags and comments – it was the first time I created a Flickr account, so it was all new to me. I thought I was the first one to digitize the report, but as it always happens when you think you are the first, you are certainly not. I found out that the report had already been digitized as part of the Historic Pittsburgh Collections. So much for being the pioneer digitizer.
Some notes and problems: What struck me when creating the master image was how big the image is when you create a really good scan (it was exceeding 100 MB’s). This is a pretty big file, and so this is indicating the space problems on your hard drive or storage unit that you encounter if you make really high quality scans. And drive space costs money. So, with limited money and hard drive space, you can definitely not save everything.
I also had some problems with writing the tags for Flickr. Flickr always combined my key words into one big word, so it looks really strange, like: 60yearsbeforethesubwaytherewasaplan. Weird, isn’t it?
Also, I wasn’t sure if it is possible to actually arrange the thumbnails where they belong (at the first page of the exhibit in Flickr). Instead, all my thumbnails are showing up as part of the images that are part of the exhibit and Flickr is creating its own thumbnails. Does anyone have any suggestions?
Reading Notes
I was roughly familiar with the story of Linux, but it was interesting to read the story how it evolved as a spin-off from UNIX, and has been developed as an open source software ever since by a large group of programmers. It was also interesting that Linux programmers became more user friendly over time. I still find it difficult to understand, though. Can you just install it on a windows machine? I browsed through the Mac OSX/Linux article (I suppose we are not supposed to read these technical pieces line by line), and what struck me was Singh's critical, but dispassionate position toward windows as a client computer and his acknowledgement that most people use windows, not the least because it is cheaper than Mac. The update on the windows roadmap reads like a promotion piece – everything is moving forward, steadily improving, if you only trust microsoft!
Overall, the readings about the different operating systems highlight that compatibility is a huge problem – compatibility between older and newer operating systems, and also compatibility between mac and windows – I am not sure about Linux. It also creates a lot of practical problems if you work at a place that has both macs and pcs – at the archives where I used to work we often had to transfer multimedia files back and forth between windows and macs, which was not always easy.
Muddiest point
The lectures were all very clear; I don’t really have a muddy point. I am just generally wondering about the implications of the constant development and updates of hardware for libraries and archives – should libraries keep versions of old hardware models (f.e. 268 PCs with 5 ¼ inch floppy drives) so that they will be able to read files on old floppies that may contain important files and re-emerge one day?
Comments
I am a little unclear on this, but are we supposed to document where we have left comments?
I commented on Kristine Harveux-Lundeen's blog:
http://2600kristineharveaux-lundeen.blogspot.com/
And on Letisha Goerner's blog:
http://letishagoerner2600.blogspot.com/
Friday, September 4, 2009
Week 1, September 1st, 2009
Muddiest Point:
One of the muddiest points of this week concerns the issue of defining “information,” “information technology,” and to distinguish this definition from “knowledge” and “wisdom” (in the DIKW pyramid). I am wondering if it may help to highlight that the definition of “information” and “information technology” depends on the context, on the way it is used, and on the history of the term. Does information really always have to be “true” or “new” as one of the definitions implied? Doesn’t it also facilitate the dissemination of “old” information? And, as Tim Notari has pointed out in his blog entry, isn’t it difficult to determine what is “true” information? And does “information” automatically lead to “knowledge”?
Readings
Lied Library @ Four Years, 2005
While not the most scintillating read, I liked the practical perspective on the experiences with maintaining the technology at Lied library and the realistic tone of the article. The article highlighted the function of the library as a gateway of providing access to computers. It also hints at the challenges and costs associated with this function, and the need to always update the technology in the library (the three year replacement cycle of pc’s). I was interested to read about the restrictions for computer usage to community users of the library (which I suppose are users who are not students), which hints at possible consequences of making the expensive technology accessible to the students – what is open to some, becomes restricted to others. I would be curious to hear how the library is doing now, and how it is dealing with budget cuts and limited resources – how can technology be maintained at this level with limited resources?
Content, nor Containers, OCLC report 2004
I think among the most interesting points of the report is the statement that libraries should help provide context for content, or information, and should help to establish authenticity and provenance of content. For example, libraries can provide information to patrons on how search engines work, which group or company developed them for what purpose, and what kind of information they include, and what information they may exclude. In this respect, reference librarians can not only help patrons to find information, but also explain different ways of accessing information. Who finances OCLC, by the way?
Clifford Lynch, Information Literacy, 1998
While more than ten years old, Clifford Lynch’s appeal to focus information technology literacy not only on the knowledge of computers, and of specific applications, but also on the infrastructure that supports the technology, and on economic, social, political, and historical problems, still resonates today. Beyond becoming familiar with specific applications, this is also what I hope to get out of this class.
One of the muddiest points of this week concerns the issue of defining “information,” “information technology,” and to distinguish this definition from “knowledge” and “wisdom” (in the DIKW pyramid). I am wondering if it may help to highlight that the definition of “information” and “information technology” depends on the context, on the way it is used, and on the history of the term. Does information really always have to be “true” or “new” as one of the definitions implied? Doesn’t it also facilitate the dissemination of “old” information? And, as Tim Notari has pointed out in his blog entry, isn’t it difficult to determine what is “true” information? And does “information” automatically lead to “knowledge”?
Readings
Lied Library @ Four Years, 2005
While not the most scintillating read, I liked the practical perspective on the experiences with maintaining the technology at Lied library and the realistic tone of the article. The article highlighted the function of the library as a gateway of providing access to computers. It also hints at the challenges and costs associated with this function, and the need to always update the technology in the library (the three year replacement cycle of pc’s). I was interested to read about the restrictions for computer usage to community users of the library (which I suppose are users who are not students), which hints at possible consequences of making the expensive technology accessible to the students – what is open to some, becomes restricted to others. I would be curious to hear how the library is doing now, and how it is dealing with budget cuts and limited resources – how can technology be maintained at this level with limited resources?
Content, nor Containers, OCLC report 2004
I think among the most interesting points of the report is the statement that libraries should help provide context for content, or information, and should help to establish authenticity and provenance of content. For example, libraries can provide information to patrons on how search engines work, which group or company developed them for what purpose, and what kind of information they include, and what information they may exclude. In this respect, reference librarians can not only help patrons to find information, but also explain different ways of accessing information. Who finances OCLC, by the way?
Clifford Lynch, Information Literacy, 1998
While more than ten years old, Clifford Lynch’s appeal to focus information technology literacy not only on the knowledge of computers, and of specific applications, but also on the infrastructure that supports the technology, and on economic, social, political, and historical problems, still resonates today. Beyond becoming familiar with specific applications, this is also what I hope to get out of this class.
Subscribe to:
Posts (Atom)