Search

Human-generated summaries are a blend of content and style, bound by the task restrictions, but are ‘subject to subjectiveness’ of the individuals summarising the documents. We study the impact of various facets that cause subjectivity such as brevity, information content and information coverage on human-authored summaries. The scale of subjectivity is quantitatively measured among various summaries using a question–answer-based cross-comprehension test. The test evaluates summaries for meaning rather than exact words based on questions, framed by the summary authors, derived from the summary. The number of questions that cannot be answered after reading the candidate summary reflects its subjectivity. The qualitative analysis of the outcome of the cross-comprehension test shows the relationship between the length of a summary, information content and nature of questions framed by the summary author.

This paper describes an approach for constructing a mixture of language models based on simple statistical notions of semantics using probabilistic models developed for information retrieval. The approach encapsulates corpus-derived semantic information and is able to model varying styles of text. Using such information, the corpus texts are clustered in an unsupervised manner and a mixture of topic-specific language models is automatically created. The principal contribution of this work is to characterise the document space resulting from information retrieval techniques and to demonstrate the approach for mixture language modelling. A comparison is made between manual and automatic clustering in order to elucidate how the global content information is expressed in the space. We also compare (in terms of association with manual clustering and language modelling accuracy) alternative term-weighting schemes and the effect of singular value decomposition dimension reduction (latent semantic analysis). Test set perplexity results using the British National Corpus indicate that the approach can improve the potential of statistical language modelling. Using an adaptive procedure, the conventional model may be tuned to track text data with a slight increase in computational cost.

Search Results

Refine search

Refine search

Actions for selected content:

2 results

On the subjectivity of human-authored summaries *

Topic-based mixture language modelling

Search Results

Refine search

Refine search

Actions for selected content:

Save Search

2 results

On the subjectivity of human-authored summaries*

Topic-based mixture language modelling

On the subjectivity of human-authored summaries *