Friday, October 22, 2010

bibtex

[1]

Anish Das Sarma, Martin Theobald, and Jennifer Widom.
Live: A lineage-supported versioned dbms.
Technical report, Stanford University, 2009.
[&nsbp;bib&nsbp;| &nsbp;http&nsbp;]

This paper presents LIVE, a complete DBMS designed for applications
with many stored derived relations, and with a need for simple versioning
capabilities when base data is modified. Target applications include,
for example, scientific data management and data integration. A key
feature of LIVE is the use of lineage (provenance) to support
modifications and versioning in this environment. In our system,
lineage significantly facilitates both: (1) efficient propagation
of modifications from base to derived data; and (2) efficient execution
of a wide class of queries over versioned, derived data. LIVE is
fully implemented; detailed experimental results are presented that
validate our techniques.


Keywords: Live



[2]

Deepavali Bhagwat, Laura Chiticariu, Wang-Chiew Tan, and Gaurav Vijayvargiya.
An annotation management system for relational databases.
In Proceedings of the Thirtieth international conference on Very
large data bases
, pages 900-911. VLDB Endowment, 2004.
bib ]

[3]

Boris Glavic and Gustavo Alonso.
Perm: Processing provenance and data on the same data model through
query rewriting.
In Proceedings of the 25th International Conference on Data
Engineering, ICDE, Shanghai, China
, pages 174-185, 2009.
bib |
DOI ]

[4]

Rajendra Bose, Robert G. Mann, and Diego Prina-Ricotti.
Astrodas: Sharing assertions across astronomy catalogues through
distributed annotation.
In IPAW, pages 193-202. Springer, 2006.
bib |
DOI ]

Keywords: AstroDas


[5]

Brooke Rhead, Donna Karolchik, Robert M. Kuhn, Angie S. Hinrichs, Ann S. Zweig,
Pauline A. Fujita, Mark Diekhans, Kayla E. Smith, Kate R. Rosenbloom,
Brian J. Raney, Andy Pohl, Michael Pheasant, Laurence R. Meyer, Katrina
Learned, Fan Hsu, Jennifer Hillman-Jackson, Rachel A. Harte, Belinda
Giardine, Timothy R. Dreszer, Hiram Clawson, Galt P. Barber, David Haussler,
and W. James Kent.
The ucsc genome browser database:update 2010.
In Nucleic acids research Journal, volume 38, pages D613-619,
2010.
bib |
DOI ]

The University of California, Santa Cruz (UCSC) Genome Browser website (http://genome.ucsc.edu/) provides a large database of publicly available sequence and annotation data along with an integrated tool set for examining and comparing the genomes of organisms, aligning sequence to genomes, and displaying and sharing users' own annotation data. As of September 2009, genomic sequence and a basic set of annotation 'tracks' are provided for 47 organisms, including 14 mammals, 10 non-mammal vertebrates, 3 invertebrate deuterostomes, 13 insects, 6 worms and a yeast. New data highlights this year include an updated human genome browser, a 44-species multiple sequence alignment track, improved variation and phenotype tracks and 16 new genome-wide ENCODE tracks. New features include drag-and-zoom navigation, a Wiki track for user-added annotations, new custom track formats for large datasets (bigBed and bigWig), a new multiple alignment output tool, links to variation and protein structure tools, in silico PCR utility enhancements, and improved track configuration tools.



[6]

Bryan C.Russell, Antonio Torralba, Kevin P.Murphy, and William T.Freeman.
Labelme: A database and web-based tool for image annotation.
International Journal of Computer Vision, 77(1-3):157-173,
2008.
bib |
DOI ]


[7]

Annie Chabert, Ed Grossman, Larry S. Jackson, Stephen R. Pietrowiz, and Chris
Seguin.
Java object-sharing in habanero.
Commun. ACM, 41(6):69-76, 1998.
bib |
DOI ]


Keywords: Habanero



[8]

Laura Chiticariu, Wang-Chiew Tan, and Gaurav Vijayvargiya.
Dbnotes: a post-it system for relational databases based on
provenance.
In Proceedings of the 2005 ACM SIGMOD international conference
on Management of data
, pages 942-944, New York, NY, USA, 2005. ACM.
bib |
DOI ]

[9]

Elizabeth F. Churchill, Jonathan Trevor, Sara Bly, Les Nelson, and Davor
Cubranic.
Anchored conversations: chatting in the context of a document.
In Proceedings of the SIGCHI conference on Human factors in
computing systems
, pages 454-461, New York, NY, USA, 2000. ACM.
bib |
DOI ]


[10]

D. Karolchik, R. Baertsch, M. Diekhans, T. S. Furey, A. Hinrichs, Y. T. Lu,
K. M. Roskin, M. Schwartz, C. W. Sugnet, D. J. Thomas, R. J. Weber,
D. Haussler, and W. J. Kent.
The UCSC Genome Browser Database.
Nucleic Acids Research, 31(1):51-54, 2003.
bib |
DOI |
arXiv |
http ]

The University of California Santa Cruz (UCSC) Genome Browser Database is an up to date source for genome sequence data integrated with a large collection of related annotations. The database is optimized to support fast interactive performance with the web-based UCSC Genome Browser, a tool built on top of the database for rapid visualization and querying of the data at many levels. The annotations for a given genome are displayed in the browser as a series of tracks aligned with the genomic sequence. Sequence data and annotations may also be viewed in a text-based tabular format or downloaded as tab-delimited flat files. The Genome Browser Database, browsing tools and downloadable data files can all be found on the UCSC Genome Bioinformatics website (http://genome.ucsc.edu), which also contains links to documentation and related technical information.





[11]

Abdulmotaleb El-Saddik, Shervin Shirmohammadi, Nicolas D. Georganas, and Ralf
Steinmetz.
Jasmine: Java application sharing in multiuser interactive
environments.
In IDMS '00: Proceedings of the 7th International Workshop on
Interactive Distributed Multimedia Systems and Telecommunication Services
,
pages 214-226, London, UK, 2000. Springer-Verlag.
bib |
DOI ]






[12]


Sean E. Ellis and Dennis P. Groth.
A collaborative annotation system for data visualization.
In AVI '04: Proceedings of the working conference on Advanced
visual interfaces
, pages 411-414, New York, NY, USA, 2004. ACM.
bib |
DOI ]


Keywords: CAV








[13]


Mohamed Y. Eltabakh, Walid G. Aref, Ahmed K. Elmagarmid, Mourad Ouzzani, and
Yasin N. Silva.
Supporting annotations on relations.
In Proceedings of the 12th International Conference on Extending
Database Technology: Advances in Database Technology
, volume 360, pages
379-390. ACM, 2009.
bib |
DOI ]







[14]


Floris Geerts, Anastasios Kementsietsidis, and Diego Milano.
imondrian: A visual tool to annotate and query scientific databases.
In International Journal of Computer Vision, 2008.
bib ]

Keywords: iMondrian








[15]


Floris Geerts, Anastasios Kementsietsidis, and Diego Milano.
Mondrian: Annotating and querying databases through colors and
blocks.
In ICDE '06: Proceedings of the 22nd International Conference on
Data Engineering
, page 82, Washington, DC, USA, 2006. IEEE Computer Society.
bib |
DOI ]

Keywords: Mondrian








[16]


Dennis P. Groth and Kristy Streefkerk.
Provenance and annotation for visual exploration systems.
IEEE Transactions on Visualization and Computer Graphics,
12(6):1500-1510, 2006.
bib |
DOI ]

Keywords: VEXS








[17]


Reid Harmon, Walter Patterson, William Ribarsky, and Jay Bolter.
The virtual annotation system.
In VRAIS '96: Proceedings of the 1996 Virtual Reality Annual
International Symposium (VRAIS 96)
, page 239, Washington, DC, USA, 1996.
IEEE Computer Society.
bib ]

Keywords: VAnno








[18]


José Kahan, Marja-Riitta Koivunen, Eric Prud'Hommeaux, and Ralph R. Swick.
Annotea: an open rdf infrastructure for shared web annotations.
Computer Networks Journal, 39:589-608, 2002.
bib |
DOI ]

Annotea is a Web-based shared annotation system based on a general-purpose
open resource description framework (RDF) infrastructure, where annotations
are modeled as a class of metadata . Annotations are viewed as statements
made by an author about a Web document. Annotations are external
to the documents and can be stored in one or more annotation servers
. One of the goals of this project has been to re-use as much existing
W3C technology as possible. We have reached it mostly by combining
RDF with XPointer, XLink, and HTTP. We have also implemented an instance
of our system using the Amaya editor/browser and a generic RDF database,
accessible through an Apache HTTP server. In this implementation,
the merging of annotations with documents takes place within the
client. The paper presents the overall design of Annotea and describes
some of the issues we have faced and how we have solved them


Keywords: Annotea








[19]


Maria M. Loughlin and John F. Hughes.
An annotation system for 3d fluid flow visualization.
In VIS '94: Proceedings of the conference on Visualization '94,
pages 273-279, Los Alamitos, CA, USA, 1994. IEEE Computer Society Press.
bib ]

Keywords: LH System








[20]


Dong Xin Luna, Halevy Alon, and Yu Cong.
Data integration with uncertainty.
The VLDB Journal, 18(2):469-500, 2009.
bib |
DOI ]







[21]


Walid G. Aref Mohamed Y. Eltabakh, Mourad Ouzzani.
bdbms - a database management system for biological data.
In CIDR 2007, Third Biennial Conference on Innovative Data
Systems Research, Asilomar, CA, USA, January 7-10, 2007, Online Proceedings
,
pages 196-206, 2007.
bib |
.pdf ]

Keywords: bdbms








[22]


Michi Mutsuzaki, Martin Theobald, Ander de Keijzer, Jennifer Widom, Parag
Agrawal, Omar Benjelloun, Anish Das Sarma, Raghotham Murthy, and Tomoe
Sugihara.
Trio-one: Layering uncertainty and lineage on a conventional dbms.
In Proceedings of CIDR conference (system demonstration), 2007.
bib |
http ]

Trio is a new kind of database system that supports data, uncertainty,
and lineage in a fully integrated manner. The first Trio prototype,
dubbed Trio-One, is built on top of a conventional DBMS using data
and query translation techniques together with a small number of
stored procedures. This paper describes Trio-One's translation scheme
and system architecture, showing how it efficiently and easily supports
the Trio data model and query language.


Keywords: Trio








[23]


Omar Benjelloun, Anish Das, Sarma Alon, and Halevy Jennifer Widom.
Databases with uncertainty and lineage.
In VLDB Journal, volume 17, pages 243-246, 2008.
bib |
DOI ]







[24]


Omar Benjelloun, Anish Das, Sarma Alon, and Halevy Jennifer Widom.
Uldbs: Databases with uncertainty and lineage.
In In VLDB, pages 953-964, 2006.
bib ]

This paper introduces uldb s, an extension of relational databases
with simple yet expressive constructs for representing and manipulating
both Lineage and Uncertainty.

Uncertain data and data lineage are two important areas of data management
that have been considered extensively but in isolation, however many
applications require the features in tandem. Fundamentally, lineage
enables simple and consistent representation of uncertain data, it
correlates uncertainty in query results with uncertainty in the input
data, and query processing with lineage and uncertainty together
presents computational benefits over treating them separately. We
show that the ULDB representation is complete, and that it permits
straightforward implementation of many relational operations. We
define two notions of ULDB minimality- data-minimal and lineage-minimal-and
study minimization of ULDB representations under both notions. With
lineage, derived relations are no longer self-contained: their uncertainty
depends on uncertainty in the base data. We provide algorithms for
the new operation of extracting a database subset in the presence
of interconnected uncertainty. Finally, we show how ULDB enable a
new approach to query processing in probabilistic databases.
s form the basis of the Trio system, under development at Stanford.








[25]


Panos K. Chrysanthis Qinglan Li, Alexandros Labrinidis.
Vip: A user-centric view-based annotation framework for scientific
data.
In Scientific and Statistical Database Management 20th
International Conference, SSDBM 2008, Hong Kong, China, July 9-11, 2008
Proceedings
. Springer Berlin / Heidelberg, 2008.
bib |
DOI ]


Keywords: ViP








[26]


Shervin Shirmohammadi, Abdulmotaleb El Saddik, Nicolas D. Georganas, and Ralf
Steinmetz.
Web-based multimedia tools for sharing educational resources.
Journal of Educational Resources in Computing (JERIC), 1, 2001.
bib |
DOI ]







[27]


S. Shirmohammadi and N. Georganas.
Jets: A java-enabled telecollaboration system.
In ICMCS '97: Proceedings of the 1997 International Conference
on Multimedia Computing and Systems
, page 541, Washington, DC, USA, 1997.
IEEE Computer Society.
bib |
DOI ]







[28]


Divesh Srivastava.
Intensional associations between data and metadata.
In Proceedings of the ACM SIGMOD International Conference on
Management of Data (SIGMOD
, pages 401-412, 2007.
bib ]

Keywords: MMS








[29]


Jenny Yuen, Bryan Russell, Ce Liu, and Antonio Torralba.
Labelme video: building a video database with human annotations.
In IEEE International Conference on Computer Vision (ICCV),
2009.
bib ]

Keywords: LabelMe video





This file was generated by
bibtex2html 1.94.

Wednesday, April 14, 2010

ORegAnno-NAR07

"ORegAnno: an open access community-driven resource for regulatory annotations"
  • an open-source, open-access database and literaturecuration system for community-based annotation of experimentallyidentified DNA regulatory regions, transcription factor binding sites and regulatory variants
  • based on open-source technology and is comprised of a MySQL database with a Java-based web application that indexes new annotations using the Lucene search engine (http://lucene.apache.org/) and provides programmatic access to the underlying data using Hibernate (http://www.hibernate.org/) and SOAP Web Services
  • this service is made for curation of publications

Tuesday, April 13, 2010

VBIGenomeACS

"Genome Annotation and Comparison System", J.Zhao, T.Xue, B.Yang, K.Williams,A.R.Wattam,R.Will,B.Sharp,R.Kenyon,O.Crasta,B.W.Sobral
  • Build with ToolBus/PathPort system which provide a generic Web Service framework to integrate backend data sources through a business logic layer
  • Web Service and XML are the core technologies for an extensible, scalable and open standard based bioinformatics tool framework
  • The system was designed to use a distributed architecture; consists of a central registration database and multiple satellite category databases
  • Central registration database provides the URL of each different category of databases as well as an entry point for each segment accession identifier. All the category databases share the same database schema, allowing for the Web service to fetch a detailed annotation from any of the databases on-the-fly.
  • Support for the annotation edition and version control was build into the database schema but the implementation of this feature is left as future work

GeneCruiser: Bioinformatics Applications Note, 2005

"GeneCruiser: a web service for the annotation of microarray data", T.Liefeld, M.Reich, J.Gould, P.Zhang, P.Tamayo, J.P.Mesirov
  • web service and web application designed to annotate genomic data
  • provides a SOAP web service interface allowing other applications to make use of its functionality

Monday, April 12, 2010

BioBuilder -annotations platform for proteins -BMC Bioinformatics 04

"BioBuilder as a database development and functional annotation platform for proteins", J.D.Navarro, N,Talreja, S.Peri, BM.Vrushabendra, BP.Rashmi,N.Padma, V.Surendranath, C.K.Jonnalagadda, PS.Kousthub, N.Deshpande, K.Shanker, A.Pandey
  • BioBuilder is a tool developed using CORBA; build as a component on top of ZOPE, an open source web applications server written in Python
  • Annotations are supervised. An approval must be granted in order for the annotation to go through
  • Made to work especially for manually creating a protein database
  • To import data from another database, the data should be in a XML format or a protein specific format

Friday, April 9, 2010

Web database query interface annotation based on user collaboration-WuhamUJ06

"Web database query interface annotation based on user collaboration", Lin Wei, Lin Can, Meng Xiaofeng; Wuahan University Journal of Natural Sciences, 2006
  • short paper for a vision based query interface annotation method
  • The GOAL of the system: to provide stimulant web service for programmatic usage of the web database
  • the interface annotation work combines automatic annotations methods and user anticipation to achieve much higher accuracy
  • the annotations refer to the "meaning" of the elements of a form
  • An URL is submitted to an URL fetcher from an interface repository; It is then transformed into a form-based query interface and an automatic process adds annotations over the block elements of the form; the annotations are send back to the user for review
  • No experiments provided; supposedly the system has 80% accuracy

Wednesday, April 7, 2010

CIDR07_bdbms-Eltabakh

"bdbms -A database management System for Biological Data", Mohamed Y.Eltabakh, Mourad Ouzzani, Walid G.Aref
  • an extensible prototype database engine for supporting and processing biological databases
  • provides framework that allows adding annotations/provenance at multiple granularities (i.e: table, tuple, column, and cell levels) archiving and restoring annotations, and querying the data based on the annotation/provenance values
  • an extension to SQL is introduced, Annotation-SQL or A-SQL, to support the processing and querying of annotations and provenance informations
  • integrates the non-traditional indexing techniques inside bdbms: (1) supporting the multidimensional datasets via multidimensional indexing techniques; (2) supporting compressed datasets via novel external-memory indexes that work over the compressed data without decompressing it
  • use of Space-partitioning trees for indexing objects in a multi-dimensional space

Monday, April 5, 2010

sigmod05_DBNotes-Tan

"DBNotes: A Post-It System for Relational Databases based on Provenance" , Laura Chiticariu, Wang-Chiew Tan, Gautav Vijayvargiya
  • This is a short presentation of the system, but more clear than the VLDB version

Wednesday, March 17, 2010

ICDE06_Mondrian-Geerts

"MONDRIAN: Annotating and querying databases through colors and blocks", Floris Geerts, Anastasios Kementsietsidis, Diego Milano
  • Use the concept of BLOCK to represent an annotated set of values and COLORS applied to the blocks to represent different annotations
  • the system can annotate both single values and the associations between multiple values
  • Does NOT account for negations. Only Select, Project, Join, Union
Observations: For an annotation mechanism to be useful in practice
  • it should be able to support the annotation of value associations
  • we should be able to query values and annotations alike (in isolations and in unison)

Tuesday, March 16, 2010

UKeScience05-AnnSciData-Bose

"Annotating scientific data:why it is important and why it is difficult", Rajendra Bose, peter Buneman, Denise Ecklund

Some existing annotation systems:
  • Swiss-Prot: annotation database produced by specialist curators
  • TrEMBL: automated annotations for proteins ( link )
  • IBMdeveloperWorks: an annotation is an XML document that is linked to a target data object
Problems and solutions:
  • "pointers" to entries in other databases are sometimes used in the fields to provide mappings between co-ordinate systems
  • it's easy to propagate the annotations through the operations of the relational algebra BUT inverting these rules is non-deterministic because the annotation in the output could have come from more than one place in the input
  • An extension of SQL can be developed to propagate the annotations. One problem is when one attribute from two different tables, with two different annotations in each table, is part of the WHERE clause but not from SELECT clause. In this case, should both annotations be sent ? The solution can be to allow user to control the flow of annotation by adding some further propagation instructions to the SQL query
  • For reasons of database security many annotations are stored externally. They require a coordinate system in order to specify how they are to be attached to the data
  • External annotations can be retrieve by given the tuple: (table_name, tuple_id, the attribute_name) -- this can be a stable coordinate system
Open problems:
  • Many existing annotation system provide only a limited ability to query over annotation values
  • One needs to know where the annotation is attached to the base data and perhaps why it is attached. How can this be captured in the database and expressed in the query ?
  • Understanding the movement of annotations in the manually curated databases

Monday, March 15, 2010

VLDBJ 05 -Annotations on Relational DBS

"An Annotation Management System for Relational Databases", Deepavali Bhagwat, Laura Chiticariu, Wang-Chiew Tan, Gaurav Vijayvargiya
  • Introduces a system to annotate the source data. They store annotations in a special attribute
  • Describe 3 methods to propagate the annotations, motivated by different needs:
    • custom
    • default scheme: if the output data is copied then the annotation is propagated
    • default-all scheme: propagate annotations according to where the data is copied from in ALL equivalent formulations of the given query
  • a new SQL language is introduced (pSQL). This extension supports PROPAGATE clause which allow a user to specify how annotations should be propagate. Does not support aggregates and bag semantics (i.e. DISTINCT keyword must be present)

Thursday, March 4, 2010

vldb09-BelieveAnn-Suciu

"Believe It or Not: Adding belief Annotations to Databases", W. Gatterbauer, M.Balazinska, N.Khoussainova, D.Suciu
  • Define a belief database as a set of belief statements. It contains base information in the form of ground tuples, annotated with belief statements
  • Describe a model of database annotations that allow users to annotate both the content and other users' annotations with belief
  • Use of a concrete semantics to annotations that helps users engage in a structured discussion on content and each other's annotations.
  • Assume, by default, that a user believes every belief statement that is in the database, unless stated otherwise
  • The annotations are of the boolean logical form: "Bob believes that Alice saw X"; "Bob does not believe that Alice saw X"
  • The experiment proved that the actual overhead of belief annotations can be significantly lower than their theoretic bound

Tuesday, March 2, 2010

Sigmod09-GrammarBasedDataCleaning-Arasu_kaushik

"A Grammar-based Entity Representation Framework for Data Cleaning", Arvind Arasu, Raghav Kaushik
  • the framework presented is a programmable module that can be used to transform "dirty" input records to one or more "clean" output records with consistent representation of entities and sub-entities
  • uses grammar based rules to decide which records are the same given the first name, last name and affiliation, knowing that all these information can be written in different ways although represent the same record.

Wednesday, February 17, 2010

ACM02DataQAssesment -Pipino

"Data Quality Assessment", L.Pipino, Y.Lee, R.Wang

Completeness can be defined as:
  • schema completeness : entities and attributes are not missing from the schema
  • column completeness : a function of missing values in a column of a table
  • population completeness : given the set of values a column should have, count how many are missing
All can be measured as a ratio =1- #incomplet items/total #of items

Monday, February 15, 2010

WWW03qDriven-Web-Zeng

"Quality driven web services composition" , L. Zeng, B. Benatallah, M. Dumas, J. Kalagnanam, Q. Sheng, WWW Conference, 2003
  • Quality criteria used: price, duration, reputation, reliability, availability
  • Composite service =an aggregation of all the component services
  • Composite service is modeled as a statechart with initial state, final state and paths to travel from the init-state to fin-state
  • price =the amount of money that a service requester has to pay for executing the operation (come from the web service providers)
  • duration =expected delay between the moment when a request is sent and the moment when the result are received; is the sum of the processing time and the transmission time
  • reliability =the probability that a request is correctly responded within a maximum expected time frame (computed from historical data)
  • availability =the probability that the service is accessible ( assumption: web services send notifications to the system about their running states )
  • reputation: the average ranking given to the service by the end users
  • MAIN IDEA: algorithms to select the optimal execution plans in terms of quality based on linear programing
  • an experiment was conducted using synthetic data.

Friday, February 12, 2010

IQ2000infoQ-Naumann

"Assesment Methods for Information Quality Criteria"
Felix Naumann, Claudia Rolker -In Proceedings of the International Conference on Information Quality, 2000
  • Previous work: IQ scores rely on questionnaires; IQ based on soudness and completeness of information sources. Algorithm provided to compute the IQ score but is still based on user inout to decide weather some information is correct or not.
  • 3 IQ classes: the user, the source and the query process
  • 3 assessment-oriented IQ criteria classes: subject-criteria scores, process-criteria scores, Object-criteria scores
COMMENTS:
  • no experiments are provided
  • User input can be seen as how many time someone accessed a particular source and how long did they stayed on that source (if this information can be extracted) or how many times user returned to that source

Thursday, February 11, 2010

VLDB99QDrivenIntegration -Naumann

"Quality-driven Integration of Heterogenous Information Systems", Felix Naumann, Ulf Leser, Johann Christoph Freytag

  • Incorporate the information quality aspect into query planning
  • Quality criteria: source-specific, query-specific and attribute-specific
  • IQ (Information Quality): is the aggregation of the scores used to rank the sources and plans
  • Three phase approach to quality-driven information integration

Friday, February 5, 2010

Journal of MIS1996BeyondAccuracy-Wang-Strong

"Beyond Accuracy: What Data Quality Means to Data Consumers" - Richard Y. Wang, Diane M. Strong ( MIT )

  • Approaches to study data quality: intuitive, theoretical, empirical
  • Attributes of the quality of data: accuracy, timeliness, precision, reliability, currency, completeness, relevancy, accessibility, interpretability
  • Overall the most important attributes: accuracy and correctness