Search results

Items from 1 to 20 out of 20 results

chapter

K nearest neighbor for text summarization using feature similarity

Taeho Jo

2017 International Conference on Communication, Control, Computing and Electronics Engineering (ICCCCEE) > 1 - 5

2017 International Conference on Communication, Control, Computing and Electronics Engineering (ICCCCEE)

In this research, we propose a particular version of KNN (K Nearest Neighbor) where the similarity between feature vectors is computed considering the similarity among attributes or features as well as one among values. The task of text summarization is viewed into the binary classification task where each paragraph or sentence is classified into the essence or non-essence, and in previous works,...

chapter

Using K Nearest Neighbors for text segmentation with feature similarity

Taeho Jo

2017 International Conference on Communication, Control, Computing and Electronics Engineering (ICCCCEE) > 1 - 5

2017 International Conference on Communication, Control, Computing and Electronics Engineering (ICCCCEE)

In this research, we propose the version of K Nearest Neighbor which considers similarity among attributes for computing the similarity between feature vectors. The text segmentation task is viewed into the binary classification where each pair of sentences or paragraphs is classified into whether we put the boundary or not, and the proposed version resulted in the successful results in previous works...

chapter

Table based KNN for categorizing words

Taeho Jo

2016 18th International Conference on Advanced Communication Technology (ICACT) > 692 - 696

2016 18th International Conference on Advanced Communication Technology (ICACT)

In this research, we propose the table based KNN as the approach to the text categorization. In previous works, we discovered that encoding texts into tables improved the performance in the text categorization, so in this research, become to consider the possibility of encoding words into tables as well as texts. In this research, we encode words into tables where entries are texts and their weights,...

chapter

Table based AHC algorithm for clustering words

Taeho Jo

2016 18th International Conference on Advanced Communication Technology (ICACT) > 570 - 575

2016 18th International Conference on Advanced Communication Technology (ICACT)

This research proposes the table based AHC algorithm as the approach to the word clustering task. The results from encoding texts into tables were successful in the previous works on the text categorization and the text clustering, and if oppositely to the case of the text encoding, texts are assumed to be elements of each word, it becomes to be possible to encode words into tables. In this research,...

chapter

Non-negative Sparse Semantic Coding for text categorization

Wenbin Zheng, Yuntao Qian

Proceedings of the 21st International Conference on Pattern Recognition (ICPR2012) > 409 - 412

2012 21st International Conference on Pattern Recognition (ICPR)

In text categorization, the dimensionality reduction methods, such as latent semantic indexing and nonnegative matrix factorization, commonly yield the dense representation that is not consistent with our common knowledge. On the other hand, the popular sparse coding methods are time-consuming and their dictionaries might contain negative entries, which is difficulty to interpret the semantic meaning...

chapter

Research on Tibetan Web Sensitive Information Detection and Warning

Xiaodong Yan, Yuan Sun, Xiaobing Zhao, Guosheng Yang

2011 7th International Conference on Wireless Communications, Networking and Mobile Computing > 1 - 4

2011 7th International Conference on Wireless Communications, Networking and Mobile Computing (WiCOM)

As an important means of monitoring public opinion through Internet or other languages media, real-time construction and supplementary of sensitive information is important for language information processing monitoring. Because the data of network has a large degree of freedom and has large information capacity and they are difficult to be controlled by the Government. In order to monitor Tibetan...

chapter

Application of an ant colony algorithm for text indexing

Nadia Lachetar, Halima Bahi

2011 International Conference on Multimedia Computing and Systems > 1 - 6

2011 International Conference on Multimedia Computing and Systems (ICMCS)

Every day, the mass of information available to us increases. This information would be irrelevant if our ability to efficiently access did not increase as well. For maximum benefit, we need tools that allow us to search, sort, index, store, and analyze the available data. We also need tools helping us to find in a reasonable time the desired information by performing certain tasks for us. One of...

chapter

Using Content and Text Classification Methods to Characterize Team Performance

K Swigger, R Brazile, G Dafoulas, F C Serce, more

2010 5th IEEE International Conference on Global Software Engineering > 192 - 200

2010 Fifth IEEE International Conference Global Software Engineering (ICGSE 2010)

Because of the critical role that communication plays in a team's ability to coordinate action, the measurement and analysis of online transcripts in order to predict team performance is becoming increasingly important in domains such as global software development. Current approaches rely on human experts to classify and compare groups according to some prescribed categories, resulting in a laborious...

chapter

A Multiclass SVM Method via Probabilistic Error-Correcting Output Codes

Zhanyi Wang, Weiran Xu, Jiani Hu, Jun Guo

2010 International Conference on Internet Technology and Applications > 1 - 4

2010 International Conference on Internet Technology and Applications (iTAP 2010)

Error-correcting output code (ECOC) is an effective approach to solve the problem of multiclass SVM. In this paper, a probabilistic approach that is based on ECOC is proposed. In the training stage, a coding scheme is predefined, and a special model is trained by samples. In the classification stage, besides the labels from SVM as usual, posterior probabilities of labels are also calculated. They...

chapter

Using semantic similarity matrix for defining operations involved in NTSO for clustering 20NewsGroups

Taeho Jo

IEEE Congress on Evolutionary Computation > 1 - 6

2010 IEEE Congress on Evolutionary Computation

In this research, we propose the similarity matrix based version of NTSO as the approach to the text clustering. For using one of traditional approaches to text clustering, documents should be encoded into numerical vectors; encoding so causes the two main problems: the huge dimensionality and the sparse distribution. In order to solve the problems, in this research, we propose to encode documents...

chapter

Study on an Improved Naive Bayesian Classifier Used in the Chinese Text Categorization

Min Zuo, Guangping Zeng, Xuyan Tu

2010 Second International Conference on Modeling, Simulation and Visualization Methods > 135 - 138

2010 Second International Conference on Modeling, Simulation and Visualization Methods (WMSVM 2010)

Text Categorization is an important research branch in the data mining domain. In this paper, an improved Naive Bayesian Classifier which is based on the Genetic Algorithms is proposed. It can make an effective Naive Bayesian classifier with excellent attributes Set in the field of text categorization. The experiments show that this method has a good classification performance.

chapter

A new feature selection algorithm in text categorization

Wei Zhao, Yafei Wang, Dan Li

2010 International Symposium on Computer, Communication, Control and Automation (3CA) > 1 > 146 - 149

2010 International Symposium on Computer, Communication, Control and Automation (3CA 2010)

A major problem with text classification problems is the high dimensionality of the feature space. This paper investigates how genetic algorithm and k-means algorithm can help select relevant features in text classification. which uses the genetic algorithm (GA) optimization features to implement global searching, and uses k-means algorithm to selection operation to control the scope of the search,...

chapter

Categorization of news articles using neural text categorizer

Taeho Jo

2009 IEEE International Conference on Fuzzy Systems > 19 - 22

2009 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE)

This research proposes the application of NTC (neural text categorizer) for categorizing news articles. Even if the research on text categorization has been progressed very much, documents should be still encoded into numerical vectors. Encoding so causes the two main problems: huge dimensionality and sparse distribution. The idea of this research as the solution to the problems is to encode documents...

chapter

Chinese Text Categorization study based on feature weight learning

Yan Zhan, Hao Chen, Su-Fang Zhang, Mei Zheng

2009 International Conference on Machine Learning and Cybernetics > 3 > 1723 - 1726

2009 Eighth International Conference on Machine Learning and Cybernetics (ICMLC)

Text categorization (TC) is an important component in many information organization and information management tasks. Two key issues in TC are feature coding and classifier design. The Euclidean distance is usually chosen as the similarity measure in K-nearest neighbor classification algorithm. All the features of each vector have different functions in describing samples. So we can decide different...

chapter

Automatic text categorization using NTC

Taeho Jo

2009 First International Conference on Networked Digital Technologies > 26 - 31

2009 First International Conference on Networked Digital Technologies (NDT 2009)

In this research, we propose NTC (Neural Text Categorizer) as the approach to text categorization. Traditional approaches to text categorization require encoding documents into numerical vectors which leads to the two main problems: huge dimensionality and sparse distribution in each numerical vector. In this research, documents are encoded into string vectors instead of numerical vectors, and a new...

chapter

Classifying Spend Descriptions with Off-the-Shelf Learning Components

S. Mukherjee, D. Fradkin, M. Roth

2008 20th IEEE International Conference on Tools with Artificial Intelligence > 1 > 53 - 60

2008 20th IEEE International Conference on Tools with Artificial Intelligence (ICTAI)

Analyzing spend transactions is essential to organizations for understanding their global procurement. Central to this analysis is the automated classification of these transactions to hierarchical commodity coding systems. Spend classification is challenging due not only to the complexities of the commodity coding systems but also because of the sparseness and quality of each individual transaction...

chapter

Research of Text Classification Technology based on Genetic Annealing Algorithm

Zhu Zhen-fang, Liu Pei-yu, Lu Ran

2008 International Symposium on Computational Intelligence and Design > 1 > 265 - 269

2008 International Symposium on Computational Intelligence and Design

Text classification has received extensive attention in recent years, which is an important means of data mining. This paper analyzed basic theory and general structure of text classification, given a text classification method based on improved genetic algorithms, introduced simulated annealing mechanism of genetic algorithm to solve the precocious easy, local optimum, and so on, using the Roocchio...

chapter

Dictionary-Based Bilingual Web Page Classification

Jicheng Liu, Chunyan Liang, Jianxun Qi

2008 4th International Conference on Wireless Communications, Networking and Mobile Computing > 1 - 4

2008 4th International Conference on Wireless Communications, Networking and Mobile Computing (WiCOM)

Web page classification poses new research challenges because of the noisy nature of the pages. For the bilingual Chinese-English web pages, it also needs to be considered that how to extract the terms of different languages exactly. A new dictionary-based multilingual text categorization approach is proposed in this paper to try to classify the Chinese-English web pages in specific domain into a...

chapter

List based matching algorithm for classifying news articles in NewsPage.com

Taeho Jo, Gwyduk Yeom

2008 IEEE International Conference on System of Systems Engineering > 1 - 5

2008 IEEE International Conference on System of Systems Engineering (SoSE)

This research proposes an alternative approach to machine learning based ones for categorizing news articles given as in plain texts. In order to use one of machine learning based approaches for the task, documents should be encoded into numerical vectors; it causes two problems: huge dimensionality and sparse distribution. The proposed approach is intended to address the two problems. In other words,...

chapter

An Improved Genetic Algorithm for Text Feature Selection

Wei Zhao, Yafei Wang

2010 International Conference on Intelligent Computing and Cognitive Informatics > 7 - 10

2010 International Conference on Intelligent Computing and Cognitive Informatics (ICICCI 2010)

High-dimensional feature space affects the quality and efficiency of text categorization. This paper investigates an improved genetic algorithm that how to help select relevant features in text classification. We follow the so-called "region growing" method to initialize the population, and uses k-means algorithm to selection operation to control the scope of the search, ensure the validity...

Filter options

Keywords:
ENCODING

Publication date

Set your own date range

Content availability

Available (19)
None (1)

Keywords

TEXT ANALYSIS (12)
CLASSIFICATION ALGORITHMS (10)
SUPPORT VECTOR MACHINES (10)
TRAINING (10)
KERNEL (6)
CLUSTERING ALGORITHMS (5)
FEATURE EXTRACTION (5)
MACHINE LEARNING (5)
SEMANTICS (5)
CLASSIFICATION (4)
GENETIC ALGORITHMS (4)
ALGORITHM DESIGN AND ANALYSIS (3)
ARTIFICIAL NEURAL NETWORKS (3)
DATA MINING (3)
DICTIONARIES (3)
GENETIC ALGORITHM (3)
GENETICS (3)
MACHINE LEARNING ALGORITHMS (3)
NATURAL LANGUAGE PROCESSING (3)
NEURAL NETS (3)
PATTERN CLASSIFICATION (3)
PATTERN CLUSTERING (3)
CHINESE TEXT CATEGORIZATION (2)
DATABASES (2)
FEATURE SELECTION (2)
FEATURE SIMILARITY (2)
HEURISTIC ALGORITHMS (2)
INDEXING (2)
K NEAREST NEIGHBOR (2)
K-MEANS ALGORITHM (2)
LEARNING (ARTIFICIAL INTELLIGENCE) (2)
NEURAL TEXT CATEGORIZER (2)
PROBABILITY (2)
SPARSE DISTRIBUTION (2)
SUPPORT VECTOR MACHINE (2)
TABLE SIMILARITY (2)
TEXT CLASSIFICATION (2)
VECTORS (2)
WEB SITES (2)
20NEWSGROUPS (1)
ACCURACY (1)
ANT COLONY ALGORITHM (1)
AUTOMATED CLASSIFICATION (1)
AUTOMATIC ENCODING DETECTION (1)
BAYESIAN METHODS (1)
BILINGUAL CHINESE-ENGLISH WEB PAGES (1)
BIOLOGICAL CELLS (1)
BMR (1)
CHEMICALS (1)
CLASSIFIER DESIGN (1)
COLLABORATION (1)
COLLABORATIVE TEAMS (1)
CONTENT ANALYSIS (1)
CONTENT CLASSIFICATION METHODS (1)
DICTIONARY-BASED BILINGUAL WEB PAGE CLASSIFICATION (1)
DICTIONARY-BASED MULTILINGUAL TEXT CATEGORIZATION (1)
DISTRIBUTED LEARNING (1)
DISTRIBUTED PROGRAMMING (1)
DOMAIN CONCEPTS (1)
DOMAIN DICTIONARY (1)
EQUATIONS (1)
ERROR CORRECTION CODES (1)
EUCLIDEAN DISTANCE (1)
FEATURE CODING (1)
FEATURE SELECTION ALGORITHM (1)
FEATURE WEIGHT (1)
FEATURE WEIGHT LEARNING ALGORITHM (1)
FEEDBACK (1)
FINITE ELEMENT METHODS (1)
GA (1)
GENETIC ANNEALING ALGORITHM (1)
GLOBAL PROCUREMENT (1)
GLOBAL SEARCHING (1)
GLOBAL SOFTWARE DEVELOPMENT (1)
GLOBAL SOFTWARE STUDENT PROJECT (1)
GROUPWARE (1)
HIERARCHICAL COMMODITY CODING SYSTEMS (1)
HIERARCHICAL TOPIC STRUCTURE (1)
HUGE DIMENSIONALITY DISTRIBUTION (1)
HUMANS (1)
IMPROVED GENETIC ALGORITHM (1)
INFORMATION MANAGEMENT (1)
INFORMATION ORGANIZATION (1)
INFORMATION RETRIEVAL (1)
INTEGRATION METHOD (1)
K-NEAREST NEIGHBOR CLASSIFICATION ALGORITHM (1)
K-NN (1)
LANGUAGE PROCESSING TOOLKITS (1)
LARGE SCALE INTEGRATION (1)
LIST BASED MATCHING ALGORITHM (1)
LOGISTIC REGRESSION (1)
LOGISTICS (1)
MACHINE LEARNING ALGORITHM (1)
MATRIX ALGEBRA (1)
MONITORING (1)
MULTICLASS SVM METHOD (1)
MULTILINGUAL PAGES (1)
NAIVE BAYES ALGORITHM (1)
more

INFONA - science communication portal

Search results

Add recipient

Sending message cancelled

Are you sure you want to cancel sending this message?

Send message

Filter options

Publication date

Date range setting

Set the date range to filter the displayed results. You can set a starting date, ending date or both. You can enter the dates manually or choose them from the calendar.

Content availability

Keywords

Reporting an error / abuse

Sending the report failed

Accessibility options